DataSet documentation claims that it Represents a potentially large set of elements. as well as Iteration happens in a streaming fashion, so the full dataset does not need to fit into memory..
I spent several hours in the official docs trying to find out how to feed a large dataset in a streaming fashion. No success. All examples use either from_tensor_slices or generator, both methods are not recommended by TensorFlow themselves because of 2GB limit for the tf.GraphDef protocol buffer and it has limited portability and scalability. It must run in the same python process that created the generator, and is still subject to the Python GIL..
The only documented way I found to make it work in a streaming fashion is to use TFRecord. It allows to stream file by file. The only huge problem: my feature input shape is (200000, 2) (all float32). To flatten and convert it to FloatList is not an option because I feed it to Conv1D with 2 parallel sequences.
It feels like either I overlooked something, or it's just poorly documented, or tf.Dataset doesn't do what it claims to do.
Is there a way to make DataSet work in a streaming fashion for a large set of elements?