How to cluster a one-dimensional dataset?

Viewed 126

How do I divide a one-dimensional dataset of integers by clusters? The picture of example data:

.

I have tried to use the methods KernelDensity and Scipy.cluster.hierarchy. Not sure if these methods fit well.

1 Answers

You can do this with something like Gaussian mixture models. Here is an example -

import numpy as np
import pandas as pd
from sklearn.mixture import GaussianMixture
%matplotlib inline

#Sample data
x = [0,200,2,1,0,1,4,4,6,14,25,43,71,93,123,194,192]
num_components = 3

#Fit a model onto the data
data = np.array(x).reshape(-1,1)
model = GaussianMixture(n_components=num_components).fit(data)

clusters = model.predict(data)
df = pd.DataFrame(list(zip(x, clusters)), columns=['data', 'clusters'])

print(df)
    data  clusters
0      0         0
1    200         1
2      2         0
3      1         0
4      0         0
5      1         0
6      4         0
7      4         0
8      6         0
9     14         2
10    25         2
11    43         2
12    71         2
13    93         2
14   123         2
15   194         1
16   192         1
Related