How to cluster large amounts of data with minimal memory usage

Viewed 733

I am using scipy.cluster.hierarchy.fclusterdata function to cluster a list of vectors (vectors with 384 components).

It works nice, but when I try to cluster large amounts of data I run out of memory and the program crashes.

How can I perform the same task without running out of memory?

My machine has 32GB RAM, Windows 10 x64, python 3.6 (64 bit)

2 Answers

You'll need to choose a different algorithm.

Hierarchical clustering needs O(n²) memory and the textbook algorithm O(n³) time. This cannot scale well to large data.

You could have a look at

However, you will have to set up some pipeline to test different numbers of clusters. It's hard to say which algorithm will work best for you, though.

Related