How do I divide a one-dimensional dataset of integers by clusters? The picture of example data:
.
I have tried to use the methods KernelDensity and Scipy.cluster.hierarchy. Not sure if these methods fit well.
How do I divide a one-dimensional dataset of integers by clusters? The picture of example data:
.
I have tried to use the methods KernelDensity and Scipy.cluster.hierarchy. Not sure if these methods fit well.
You can do this with something like Gaussian mixture models. Here is an example -
import numpy as np
import pandas as pd
from sklearn.mixture import GaussianMixture
%matplotlib inline
#Sample data
x = [0,200,2,1,0,1,4,4,6,14,25,43,71,93,123,194,192]
num_components = 3
#Fit a model onto the data
data = np.array(x).reshape(-1,1)
model = GaussianMixture(n_components=num_components).fit(data)
clusters = model.predict(data)
df = pd.DataFrame(list(zip(x, clusters)), columns=['data', 'clusters'])
print(df)
data clusters
0 0 0
1 200 1
2 2 0
3 1 0
4 0 0
5 1 0
6 4 0
7 4 0
8 6 0
9 14 2
10 25 2
11 43 2
12 71 2
13 93 2
14 123 2
15 194 1
16 192 1