Cluster features based on their attributes

Viewed 427

I have a set of 5000 points each has an attribute that is set to a value from 0 to 5. I am trying to cluster this data to reduce the amount of points drawn but wish to create clusters containing only the features that have the same attribute value. Having searched around I discovered this extended clustering example from Openlayers 2.

http://dev.openlayers.org/examples/strategy-cluster-extended.html

However I see no information on how to implement this in Openlayers 3 and above? Is there something simple I am missing in order to achieve this?

Many thanks for the help on this and any advice would be much appreciated!

2 Answers

You would need a cluster source and layer for each value. You could use the geometry function as a filter

cluster0 = new Cluster({
  source: vectorSource,
  geometryFunction: function(feature) {
    if (feature.get('attribute') == '0') {
      return feature.getGeometry();
    }
    return null;
  }
});

cluster1 = new Cluster({
  source: vectorSource,
  geometryFunction: function(feature) {
    if (feature.get('attribute') == '1') {
      return feature.getGeometry();
    }
    return null;
  }
});

This doesn't really sound like a clustering question. It seems like you just need to group by value 1, 2, 3, 4, and 5. If you really want to do clustering, you can do it like this.

import pandas as pd
import numpy as np
from numpy import vstack,array
from numpy.random import rand
from scipy.cluster.vq import kmeans,vq
from sklearn.cluster import KMeans

df= pd.read_csv("filename.csv") 

#format the data as a numpy array to feed into the K-Means algorithm
data = np.asarray([np.asarray(df['Value1']),np.asarray(df['Value2'])])

X = data
distorsions = []
for k in range(2, 20):
    k_means = KMeans(n_clusters=k)
    k_means.fit(X)
    distorsions.append(k_means.inertia_)

fig = plt.figure(figsize=(15, 5))
plt.plot(range(2, 20), distorsions)
plt.grid(True)
plt.title('Elbow curve')

# computing K-Means with K = 5 (5 clusters)
centroids,_ = kmeans(data,5)
# assign each sample to a cluster
idx,_ = vq(data,centroids)

details = [(name,cluster) for name, cluster in zip(df.index,idx)]

for detail in details:
    print(detail)
Related