How to use sklearn's IncrementalPCA partial_fit

Viewed 4334

I've got a rather large dataset that I would like to decompose but is too big to load into memory. Researching my options, it seems that sklearn's IncrementalPCA is a good choice, but I can't quite figure out how to make it work.

I can load in the data just fine:

f = h5py.File('my_big_data.h5')
features = f['data']

And from this example, it seems I need to decide what size chunks I want to read from it:

num_rows = data.shape[0]     # total number of rows in data
chunk_size = 10              # how many rows at a time to feed ipca

Then I can create my IncrementalPCA, stream the data chunk-by-chunk, and partially fit it (also from the example above):

ipca = IncrementalPCA(n_components=2)
for i in range(0, num_rows//chunk_size):
    ipca.partial_fit(features[i*chunk_size : (i+1)*chunk_size])

This all goes without error, but I'm not sure what to do next. How do I actually do the dimension reduction and get a new numpy array I can manipulate further and save?

EDIT
The code above was for testing on a smaller subset of my data – as @ImanolLuengo correctly points out, it would be way better to use a larger number of dimensions and chunk size in the final code.

1 Answers
Related