I have a ~30GB (~1.7 GB compressed | 180K rows x 32K columns) matrix saved in a csv format. I would like to convert this matrix to sparse format to be able to load the full dataset in memory for machine learning with sklearn. The cells that are populated contain float numbers less than 1. A caveat of the large matrix is the target variable is stored as the last column. What is the best method to allow this large matrix to be utilized in sklearn? I.E. How can you transition the ~30GB csv into a scipy sparse format without loading the original matrix into memory?
Pseudocode
- Remove target variable (keep order intact)
- Convert ~30 GB matrix to sparse format (Help!!)
- Load sparse format into memory and target variable to run machine learning pipeline (How would I do this?)