python code to find feature importances after kmeans clustering

Viewed 7194

I researched the ways to find the feature importances (my dataset just has 9 features).Following are the two methods to do so, But i am having difficulty to write the python code.

I am looking to rank each of the features who's influencing the cluster formation.

  1. Calculate the variance of the centroids for every dimension. The dimensions with the highest variance are most important in distinguishing the clusters.

  2. If you have only a little number of variables you could do some kind of leaving-one-out test (remove 1 var and redo clustering). Also keep in mind that k-means depends on the initialization, so you want to keep that fixed when you redo the clustering.

Any python codes to accomplish this?

3 Answers

Lets say we have X with 200 samples and 9 variables, and synthetically make them have two clusters, and just for the sake of visualization, we fill two of these variables at each time.

import numpy as np
import matplotlib.pyplot as plt
import sklearn
X = np.zeros((200,4))

Feature1_1 = np.random.normal(loc=40, scale=1.0, size=100)
Feature1_2 = np.random.normal(loc=70, scale=3.0, size=100)

Feature2_1 = np.random.normal(loc=20, scale=4.0, size=100)
Feature2_2 = np.random.normal(loc=50, scale=1.0, size=100)

X[:100,0]=Feature1_1
X[100:,0]=Feature1_2
X[:100,1]=Feature2_1
X[100:,1]=Feature2_2

plt.figure(figsize = (5,5))
plt.scatter(X[:,0],X[:,1])
plt.grid()
plt.xlabel('Feature 2',fontsize=18)
plt.ylabel('Feature 1',fontsize=18)

enter image description here

Now, lets populate a new feature with a higher variance.

Feature3_1 = np.random.normal(loc=40, scale=300.0, size=100)
Feature3_2 = np.random.normal(loc=43, scale=280.0, size=100)

Feature2_1 = np.random.normal(loc=20, scale=4.0, size=100)
Feature2_2 = np.random.normal(loc=50, scale=1.0, size=100)


X[:100,2]=Feature3_1
X[100:,2]=Feature3_2

X[:100,1]=Feature2_1
X[100:,1]=Feature2_2

plt.figure(figsize = (5,5))
plt.scatter(X[:,2],X[:,1])
plt.grid()
plt.xlabel('Feature 3',fontsize=18)
plt.ylabel('Feature 2',fontsize=18)

enter image description here

and a last one with higher variance too

Feature3_1 = np.random.normal(loc=40, scale=300.0, size=100)
Feature3_2 = np.random.normal(loc=43, scale=280.0, size=100)

Feature4_1 = np.random.normal(loc=20, scale=40.0, size=100)
Feature4_2 = np.random.normal(loc=22, scale=40.0, size=100)


X[:100,2]=Feature3_1
X[100:,2]=Feature3_2

X[:100,3]=Feature4_1
X[100:,3]=Feature4_2

plt.figure(figsize = (5,5))
plt.scatter(X[:,2],X[:,3])
plt.grid()
plt.xlabel('Feature 3',fontsize=18)
plt.ylabel('Feature 4',fontsize=18)

enter image description here

now, lets cluterize them with k-means

from sklearn.cluster import KMeans
kmeans = KMeans(n_clusters=2, random_state=0).fit(X)

now, lets visualize those clusters.

f1=0
f2=1

plt.figure(figsize = (5,5))
plt.scatter(X[kmeans.labels_==0][:,f1],X[kmeans.labels_==0][:,f2])
plt.scatter(X[kmeans.labels_==1][:,f1],X[kmeans.labels_==1][:,f2])
plt.grid()
plt.xlabel('Feature 3',fontsize=18)
plt.ylabel('Feature 2',fontsize=18)

enter image description here

f1=2
f2=1

plt.figure(figsize = (5,5))
plt.scatter(X[kmeans.labels_==0][:,f1],X[kmeans.labels_==0][:,f2])
plt.scatter(X[kmeans.labels_==1][:,f1],X[kmeans.labels_==1][:,f2])
plt.grid()
plt.xlabel('Feature 3',fontsize=18)
plt.ylabel('Feature 2',fontsize=18)

enter image description here

f1=2
f2=3

plt.figure(figsize = (5,5))
plt.scatter(X[kmeans.labels_==0][:,f1],X[kmeans.labels_==0][:,f2])
plt.scatter(X[kmeans.labels_==1][:,f1],X[kmeans.labels_==1][:,f2])
plt.grid()
plt.xlabel('Feature 3',fontsize=18)
plt.ylabel('Feature 2',fontsize=18)

enter image description here

we can now very clearly see that feature 3 and 4 are the only important ones. Note that normalized features would lead to a completely different result.

And finally, we automate this with:

for feature in range(X.shape[1]):
    mean1 = X[kmeans.labels_==0][:,feature].mean()
    mean2 = X[kmeans.labels_==1][:,feature].mean()
    
    var1 = X[kmeans.labels_==0][:,feature].var()
    var2 = X[kmeans.labels_==1][:,feature].var()
    
    print('feature:',feature,'Mean difference:',round(abs(mean1-mean2),3),'Total Variance:',round((var1+var2),3))

resulting in:

feature: 0 Mean difference: 1.69 Total Variance: 459.464 
feature: 1 Mean difference: 0.879 Total Variance: 449.829 
feature: 2 Mean difference: 66.213 Total Variance: 154932.184 
feature: 3 Mean difference: 2.076 Total Variance: 2731.953

The comment is clearly a copy and paste of the documentation, but it wasn't listened the need of the question. these solution are coming from the supervised learning section of the scikit learn user guide.

enter image description here

Consider doing feature selection like this.

import pandas as pd
import numpy as np
import seaborn as sns
from sklearn.feature_selection import SelectKBest
from sklearn.feature_selection import chi2

# UNIVARIATE SELECTION

data = pd.read_csv('C:\\Users\\Excel\\Desktop\\Briefcase\\PDFs\\1-ALL PYTHON & R CODE SAMPLES\\Feature Selection - Machine Learning\\train.csv')
X = data.iloc[:,0:20]  #independent columns
y = data.iloc[:,-1]    #target column i.e price range

#apply SelectKBest class to extract top 10 best features
bestfeatures = SelectKBest(score_func=chi2, k=10)
fit = bestfeatures.fit(X,y)
dfscores = pd.DataFrame(fit.scores_)
dfcolumns = pd.DataFrame(X.columns)
#concat two dataframes for better visualization 
featureScores = pd.concat([dfcolumns,dfscores],axis=1)
featureScores.columns = ['Specs','Score']  #naming the dataframe columns
print(featureScores.nlargest(10,'Score'))  #print 10 best features


# FEATURE IMPORTANCE
data = pd.read_csv('C:\\your_path\\train.csv')
X = data.iloc[:,0:20]  #independent columns
y = data.iloc[:,-1]    #target column i.e price range
from sklearn.ensemble import ExtraTreesClassifier
import matplotlib.pyplot as plt
model = ExtraTreesClassifier()
model.fit(X,y)
print(model.feature_importances_) #use inbuilt class feature_importances of tree based classifiers
#plot graph of feature importances for better visualization
feat_importances = pd.Series(model.feature_importances_, index=X.columns)
feat_importances.nlargest(10).plot(kind='barh')
plt.show()

enter image description here

# Correlation Matrix with Heatmap
data = pd.read_csv('C:\\your_path\\train.csv')
X = data.iloc[:,0:20]  #independent columns
y = data.iloc[:,-1]    #target column i.e price range
#get correlations of each features in dataset
corrmat = data.corr()
top_corr_features = corrmat.index
plt.figure(figsize=(20,20))
#plot heat map
g=sns.heatmap(data[top_corr_features].corr(),annot=True,cmap="RdYlGn")

enter image description here

Dataset is available here:

https://www.kaggle.com/iabhishekofficial/mobile-price-classification#train.csv
Related