I want to find optimal k for using k-means on a data set. I use code below:
Sum_of_squared_distances = []
for k in range(1,15):
print(k)
Sum_of_squared_distances.append(KMeans(n_clusters=k).fit(x).inertia_)
plt.plot(range(1,15), Sum_of_squared_distances, 'bx-')
plt.xlabel('k')
plt.xticks(range(1,15))
plt.ylabel('Sum_of_squared_distances')
plt.title('Elbow Method For Optimal k')
plt.savefig('optimal-k.jpg')
plt.show()
as you can see for some values of k, higher k has worse result. I want to know is this possible or I am doing something wrong? some in depth explanation would be appreciated.
