I have a CSV file with the shape of (9177254, 7). It takes 500 MB of disk space. When I import the csv as an array, it takes 5GB of ram! When I import it as a dataframe, it takes about 1GB.
Each element of my training set (dp) should be an "array" in the shape of 10000x7. The 1-100000 row is the datapoint 1, the 2-10001 row is the datapoint 2, etc, with a total number of datapoint of 9167254
If I have unlimited memory, I can generate the "training_set" array (9167254x10000x7) which will take about 10TB of memory (which is unpractical).
model.compile(optimizer='adam',
loss=tf.keras.losses.SparseCategoricalCrossentropy(from_logits=True),
metrics=['accuracy'])
history = model.fit(x_train, y_train, epochs=100,
validation_data=(x_test, y_test))
The training input "x_test" must be an array to my knowledge. Is it possible to find a way around? Can I define x_test as a mapping (or function) of dataframe to save memory space?