How do you deal with very large dataset when creating the matrix for recommender system?

Viewed 224

I am trying to create a transaction and product groups matrix but I have a very large transaction data (over 10,000,000 rows) and around 100 product groups. When I try to create a pivot table using this code

df.pivot(index='transaction_id', columns='product_group', values='ratings')

It returned values error "Unstacked DataFrame is too big, causing int32 overflow"

Is there anyway to deal with this issue other than decrease the size of the data?

Thanks!

1 Answers

Convert your axes columns to categories:

df['transaction_id'] = df['transaction_id'].astype('category')
df['product_group'] = df['product_group'].astype('category')

Make a sparse matrix using the encodings:

arr = csr_matrix((df['ratings'].values, (df['transaction_id'].cat.codes, df['product_group'].cat.codes)))

Then you just have to keep track of the order of your axes (df['transaction_id'].cat.categories will give you the labels that should be applied to the rows for example).

Related