I implement a regression model of a multivariate time series in TensorFlow. The number of input features is 9 and the length of the sequences is variable. Lets say the sequence length is between 32 and 512, but most of them are nearer towards the lower end.
As I understand, to train in a batched manner, all samples need to have the same length. If I choose the maximum length of all samples, the data tensor will have shape=[batch_size, 512, 9]. To achieve this, I need to zero-pad each sample with length<512.
My model consists of some simple convolutions plus maxpool-layers:
class MyModel(tf.keras.Model):
def __init__(self, params):
super(MyModel, self).__init__()
self.params = params
def build(self, inputs_shape):
self.conv1 = tf.keras.layers.Conv1D(filters=32, kernel_size=3, activation='relu', input_shape=inputs_shape[1:])
self.pool1 = tf.keras.layers.MaxPool1D(pool_size=2)
self.conv2 = tf.keras.layers.Conv1D(filters=64, kernel_size=3, activation='relu')
self.pool2 = tf.keras.layers.MaxPool1D(pool_size=2)
self.conv3 = tf.keras.layers.Conv1D(filters=64, kernel_size=3, activation='relu')
self.gpool1 = tf.keras.layers.GlobalAveragePooling1D()
self.dense2 = tf.keras.layers.Dense(32, activation='relu')
self.dense3 = tf.keras.layers.Dense(1)
def call(self, x):
...
So I train my model with a sequence length of 512. In my test set I only have a maximum sequence length of lets say 128. So I only need to pad less.
Does the different input sequence length have an impact on the outcome? Since when the majority of samples for an individual sequence is padded, the majority of the output of the convolutional layer will be zero. This means, that the input to the global average pooling will consist of a lot of zeros (dependent on how many samples were padded). When building the average, the presence of many zeros will change the overall outcome.
This means, that differnet lengths of input samples will have differnet output.
So my question is:
Is there any way to circumvent this problem? Either by A: masking the padded samples for later correction? Or B: training my model with a batch size of 1 so I do not need to pad any sample. --> Is this even possible since the input tensor to the model is different for each training step.
As far as I understand, this should not be the case when using global maximum pooling, since the additional zeros after the convolkution should not have an impact. However, my experiments show, that there is indeed a difference based on the padding size.
Now, I´m not sure, if my thoughts are correct, or am I missing something?
For inference it also might happen that a never seen sample is even longer than 512 time stamps. This would mean that I need to feed the sample with it´s original dimensions to the model.
The concept of Global average pooling is used in many differnet models and I want to implement the TimeInception network which uses this principle as well.