Adding new document into existing cluster

Viewed 239

I am doing clustering(kmeans) for a large dataset. Now wanted to add new data into existing cluster.

Here is my idea:

  1. Calculate the euclidian distance of a new data point from all centroids and find the minimum of those distances.

  2. Check if the minimum distance is less than the threshold value. If true, we assign the new data point to the corresponding cluster. Then, update the cluster center of that cluster.

  3. If False, create a new cluster and assign the new data point as its center. Also, the data point becomes a part of the cluster.

In step 2, what will be the threshold value that I should use. Please share your idea.

I am thinking, by calculating the intracluster distance from each cluster and take the maximum distance from them would be my threshold value.

I am following the article here

1 Answers

Instead of threshold, can't you use an internal validation such as silhouette score to see whether you need to add the number of clusters by one or just fit the new data point into one of the existing clusters?

And, regarding your suggestion about threshold, let's say you have two clusters C1 and C2 far from each other (let's say the distance between their centers are 10), and the distance between their centers and their furthest member are 1 and 1.1. Now, you have a new point whose distance from the (updated or original) centers of C1 is 1.2. What is your call? Since it is slightly bigger 1, but at the same time bigger than 1.1, you just put it into a new cluster (?!). As you can see, that's not a reasonable approach.

If you insist on using the threshold, here is an idea: You can find the distance of the new point to the closest center (call it d1) and the next closest center (call it d2). If, for instance, d1/d2 is less than 0.5 (a threshold), you can say the new point belongs to the closest group, if not, it means you cannot determine it belongs to which group. So, you then create a new cluster

Related