Feature importance in regression models

Viewed 126

I used KNN, Decision Tree, Random Forest and ANN to make predictions on my data using Python I have 9 predictors. The question I'm having is which of them are not contributing. Decision Tree, Random Forest allow to run the feature importance. I did so and it indicated that that 3 predictors contribute very little. So it seems i can delete them from the dataset. For KNN and ANN no model.feature_importances_

Would it be correct to assume that for KNN and ANN the same predictors also don't contribute? Or does Feature importance depend on the model (i.e for KNN for example those will be different than the ones for Random forest)

Thank you

1 Answers

I concur with Ben. Here is a generic example of using a Random Forest Regressor to find the importance of each feature in the data set.

from sklearn.datasets import load_boston
from sklearn.ensemble import RandomForestRegressor
import numpy as np
import pandas as pd
import matplotlib.pyplot as plt

#Load boston housing dataset as an example
boston = load_boston()


X = boston["data"]
Y = boston["target"]
names = boston["feature_names"]
reg = RandomForestRegressor()
reg.fit(X, Y)
print("Features sorted by their score:")
print(sorted(zip(map(lambda x: round(x, 4), reg.feature_importances_), names), 
             reverse=True))


boston_pd = pd.DataFrame(boston.data)
print(boston_pd.head())

boston_pd.columns = boston.feature_names
print(boston_pd.head())

# correlations


features = boston.feature_names
importances = reg.feature_importances_
indices = np.argsort(importances)

plt.title('Feature Importances')
plt.barh(range(len(indices)), importances[indices], color='#8f63f4', align='center')
plt.yticks(range(len(indices)), features[indices])
plt.xlabel('Relative Importance')
plt.show()

Result:

enter image description here

Related