How can I improve the silhouette score of my k-?means clustering

Viewed 90

I have a dataset with 18000 lines about some Customers, like this: enter image description here

and I am trying to do some clustering using k-means algorithm. Since I have both categorical and continuous variables I created some dummies for the categorical variables

#dummy codification
dataML=pd.get_dummies(dataML)

#print(dataML.head())
X=dataML
mms=StandardScaler()
Xnorm=mms.fit_transform(X)

I then proceed to do the clustering

#Fit k-means , k=3
#3 clusters, 10 initializations(find 10 times the initial clusters(random), max iterations, seed)
km=KMeans(n_clusters=4,n_init=10,max_iter=30,random_state=42)
y_kmeans=km.fit_predict(Xnorm)

#K-labels assigned
print("Labels assigned: ")
print(y_kmeans)

#The lowest SSE value
print("The lowest SSE value: " ,km.inertia_)

#The number of iterations required to converge
print("Num iterations to converge: ",km.n_iter_)

print("Final centers")
print(km.cluster_centers_)

#Clustering evaluation
#Silhouette score

#the closest to 1 the better
silSc=silhouette_score(X,y_kmeans,metric="euclidean")
print("Silhouette score: " , round(silSc,3))

But I get a negative silhouette score value. Is there something wrong with my code?

I tried removing the StandardScaler and the silhoutte score went up to 0,6. Why does this happen?

0 Answers
Related