I'm working on a deep learning project with about 700GB of table-like time series data in thousands of .csv files (each about 15MB).
All the data is on S3 and it needs some preprocessing before being fed into the model. The question is how to best go about automating the process of loading, preprocessing and training.
Is a custom keras generator with some built in preprocessing the best solution?