"Generate" a training set from a dataframe without consuming more memory

Viewed 24

I have a CSV file with the shape of (9177254, 7). It takes 500 MB of disk space. When I import the csv as an array, it takes 5GB of ram! When I import it as a dataframe, it takes about 1GB.

Each element of my training set (dp) should be an "array" in the shape of 10000x7. The 1-100000 row is the datapoint 1, the 2-10001 row is the datapoint 2, etc, with a total number of datapoint of 9167254

If I have unlimited memory, I can generate the "training_set" array (9167254x10000x7) which will take about 10TB of memory (which is unpractical).

model.compile(optimizer='adam',
              loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True),
              metrics=['accuracy'])

history = model.fit(x_train, y_train, epochs=100, 
                    validation_data=(x_test, y_test))

The training input "x_test" must be an array to my knowledge. Is it possible to find a way around? Can I define x_test as a mapping (or function) of dataframe to save memory space?

0 Answers
Related