how to visualize multi-dimensionnal clusters in Python?

Viewed 558

I am trying to test 3 algorithms of clustering (K-means , SpectralClustering ,Mean Shift) in Python. I have a datset containing 26 columns and several thousand rows ,i need some help with a high dimensional data-set (subset is shown below).

UserID  Communication_dur   Lifestyle_dur   Music & Audio_dur   Others_dur  Personnalisation_dur    Phone_and_SMS_dur   Photography_dur Productivity_dur    Social_Media_dur    System_tools_dur    ... Music & Audio_Freq  Others_Freq Personnalisation_Freq   Phone_and_SMS_Freq  Photography_Freq    Productivity_Freq   Social_Media_Freq   System_tools_Freq   Video players & Editors_Freq    Weather_Freq
1   63  219 9   10  99  42  36  30  76  20  ... 2   1   11  5   3   3   9   1   4   8
2   9   0   0   6   78  0   32  4   15  3   ... 0   2   4   0   2   1   2   1   0   0

I have to cluster data with very high dimensions. I want to know how it can be achieved accurately as possible. How can I visualize the clusters and data points?

P.S: after some search I have realised that one can apply PCA for dimension reduction but I want to know how it can be used .

2 Answers

PCA, t-SNE, and UMAP are tools that might help you to achieve a good visualization. Just google PCA sklearn and read some examples. You can reduce the dimensions of your data (let's call it your new features) to two or three. Then, dedicate a particular color to members of each cluster. Your plot should show data points (with same color) close together as one cluster. However, be aware, that using such tools don't necessarily keep the relative distances the data points have with each other in their original space. In other words, your clustering process might perform well, but such "very good" performance may not be "obvious" in 2 or 3 dimension.

Here is a very short example I got from Scikit-Learn webpage:

from sklearn import decomposition
pca = decomposition.PCA(n_components=2)
pca.fit(X) 
X = pca.transform(X)

In this example, X is the feature matrix. In this example, n_components=2 means that you want to reduce your data into a new 2-dimension space. After .tranform(), your new feature matrix X has the same number of observations as its row but only three column rather than all of its previous features.

Then, after clustering, allocated different colors to different clusters and plot them in 2D space. you can choose different n_components but as you increase it, it is getting harder to visualize it. For instance, for n_components=3, you can use 3D. For, 4D you may use color to show the extra dimension. For 5D, you may want to change sizes of data points as a way to show the new dimension.

I attached a picture where you can see how you can use visualization based on PCA to show the performance of a classification model on train & test set. You can just use it for clustering by allocation one color to each group. enter image description here

Related