Kmean on multivariate normal distribution samples

Viewed 121

I am new to Python. I want to perform K-means clustering on samples generated by two multivariate normal distributions. I generated the samples and performed the K-means clustering but I get an error when I want to plot the clusters. There should be some dimensional problem that I am missing. please see my code below:

mean1 = [-1, -1.5]
cov1 = [[1, .2], [.2, 1]]
x1, y1 = np.random.default_rng().multivariate_normal(mean1, cov1, 100).T
mean2 = [1, 1.5]
cov2 = [[2, .1], [.1,2]]
x2, y2 = np.random.default_rng().multivariate_normal(mean2, cov2, 100).T
X = np.mat([x1,x2]).reshape(-1).transpose()
Y= np.mat([y1,y2]).reshape(-1).transpose()
kmeans = KMeans(n_clusters=2)
kmeans.fit(X)
y_kmeans = np.array(kmeans.predict(X))
plt.scatter(X[:, 0], X[:, 1], c=y_kmeans, s=50, cmap='viridis')
1 Answers

Since it is not that clear, these are the two cases. If you are trying to generate an X where each item has 2 values, x1 and x2, then you are concatenating them wrong. You are generating an X of shape (200, 1) instead of (100, 2). For that, you can use

X = np.stack((x1, x2)).reshape(-1, 2) # (100, 2)
or
X = np.concatenate((x1[:, None], x2[:, None]), axis=-1)

And for Y

Y = np.stack((y1, y2)).reshape(-1) # (200,)
or
Y = np.concatenate((y1, y2))

Afterwards, you can get your plot with plt.scatter(X[:, 0], X[:, 1], c=y_kmeans, s=50, cmap='viridis') without erros.

clustering on 2D

If you are trying to cluster samples that lie on a line, i.e, your X is a 1D array, your plot fails because X contains the samples and your y values should simply be an array of zeros. In that case, your Y is the same as before but

X = np.stack((x1, x2)).reshape(-1, 1) # (200, 1)
or 
X = np.concatenate((x1, x2), axis=0).reshape(-1, 1)
# The reshape(-1, 1) is necessary for sklearn's KMeans if the data has a single feature

Then, if you plot your result with plt.scatter(X, np.zeros((X.shape[0],)), c=y_kmeans, s=50, cmap='viridis'), you get

clustering in 1D

Related