From my understanding of K-means clustering, k is chosen, centroid locations are selected, samples are assigned, and then the centroid moves to the mean of the samples, until there is no more movement.
So in this example, if 3 centroids are chosen with locations of (5,1), (5,10), and (5,20), what happens? I would expect all the samples to be assigned to the (5,1) centroid which would then move to the mean of the data (5,0) and the algorithm would terminate with all of the samples belonging to 1 cluster (the other centroids not moving, and with 1 iteration).
However, when I implement this in Python, using KMeans, it doesn't behave like this:
In [3]: xa = [1,2,3,4,5,6,7,8,9,10]
...: xb = [0,0,0,0,0,0,0,0,0,0]
...:
...: data = pd.DataFrame({"Xa": xa, "Xb":xb}, columns=["Xa", "Xb"])
...:
...: initial = np.array([[5,1], [5,10], [5,20]])
...:
...: km = KMeans(n_clusters=3,init=initial, n_init=1, max_iter=1).fit(data)
...: print(km.cluster_centers_)
[[ 4.5 0. ]
[10. 0. ]
[ 9. 0. ]]
Letting the algorithm run with no limit on the number of iterations results in the following, which also doesn't make sense to me:
In [4]: km = KMeans(n_clusters=3,init=initial, n_init=1).fit(data)
...: print(km.cluster_centers_)
[[3. 0. ]
[9.5 0. ]
[7. 0. ]]
Could anyone explain my misunderstanding here?
