E.g. imagine I use the Librispeech dataset via TFDS (or whatever dataset, including sequences of varying length of data), and then use padded_batch to create batches, e.g. like this:
import tensorflow_datasets as tfds
dataset = tfds.load(name="librispeech", split="train_clean100")
dataset = dataset.shuffle(1024)
dataset = dataset.padded_batch(32)
Now when iterating through the resulting dataset, i.e. over the (padded) batches, how would I know the original sequence lengths in the padded batch? Or is this information lost at this point? How would I extend the pipeline to include it? Is there a special dataset like AddSeqLengthInfoDataset or so? This would need to run before the padded_batch, right?
(This is basically an equivalent of my question for TF PaddingFIFOQueue but for tf.data.Dataset.)
Is there some example? (I wonder a bit that I have not found anything about this. I would assume this is a pretty standard requirement when you work on sequences, that you need to have the information about the original sequence lengths, or not?)