I am new with Tensorflow input data pipeline development. I have 40 csv files each of which is 100GB and has some of my input features for training a classifier. I want to use python and tensorflow-gpu to train a classifier. Of course, I can not merge all 40 csv files and then read it into RAM because I do not have 4TB RAM (40x100GB=4TB RAM!!). It also makes no sense to read all 40 csv files every time, select some rows, and then merge selected rows because reading these very large files million times during training will slow down my input data pipeline. Before training any model, I am thinking of splitting each of these 40 csv files into 10000 parts, merging corresponding parts in all 40 csv files (SO I will have 10000 csv files instead), then using their address and class label in the following code:
filenames = ["filename_00001.csv", "filename_00002.csv", ...,
"filename_9999.csv", "filename_10000.csv"]
labels = [1, 1, ..., 0, 0...]
dataset = tf.data.Dataset.from_tensor_slices((filenames, labels))
dataset = dataset.shuffle(buffer_size=len(filenames))
dataset = dataset.map(...) # Read selected subset of 10000 csv files, preprocess, repeat, batch...
Is there a better way to read this data into 100GB RAM faster? Also, I have several million features. Can I feed all features into a fully connected model or will I encounter some practical issues because my GPUs have limited RAM?