I'm trying to get the feature importances for a Regression model. I have 58 independent variables and one dependent variables. Most of the independent variables are numerical and some are binary.
First I used this:
X = dataset.drop(['y'], axis=1)
y = dataset[['y']]
# define the model
model = LinearRegression()
# fit the model
model.fit(X, y)
# get importance
importance = model.coef_[0]
print(model.coef_)
print(importance)
# summarize feature importance
for i,v in enumerate(importance):
print('Feature: %0d, Score: %.5f' % (i,v))
# plot feature importance
pyplot.bar([x for x in range(len(importance))], importance)
pyplot.show()
and got the following results: Feature Importance Plot
Then I used MinMaxScaler() to scale the data before fitting the model:
scaler = MinMaxScaler()
dataset[dataset.columns] = scaler.fit_transform(dataset[dataset.columns])
print(dataset)
X = dataset.drop(['y'], axis=1)
y = dataset[['y']]
# define the model
model = LinearRegression()
# fit the model
model.fit(X, y)
# get importance
importance = model.coef_[0]
print(model.coef_)
print(importance)
# summarize feature importance
for i,v in enumerate(importance):
print('Feature: %0d, Score: %.5f' % (i,v))
# plot feature importance
pyplot.bar([x for x in range(len(importance))], importance)
pyplot.show()
which led to the following plot: Feature Importance Plot after using MinMaxScaler
As you can see in the upper left corner it is 1e11, which means the largest values are negative 60 billion. What am I doing wrong here? And is it even the right approach to use MinMaxScaler?