I have a data frame that contains 27k records as a training set and another testing dataset with 4k records. Both datasets have 25 features each.
x_train shape: (27000, 25),
x_test shape: (4000, 25)
Example of data in the training set:
|Subject ID|Feat_1|Feat_2|Feat_X|Hr_count|Label|
|s0001 | 89| 31 | 43 | 1 | 0 |
|s0001 | 94| 32 | 68 | 2 | 0 |
|s0001 | 38| 90 | 86 | 3 | 0 |
|s0001 | 79| 34 | 78 | 4 | 1 |
|s0001 | 85| 24 | 70 | 5 | 1 |
|s0002 | 7 | 9 | 32 | 1 | 0 |
|s0002 | 60| 56 | 72 | 2 | 0 |
|s0002 | 68| 72 | 23 | 3 | 0 |
|s0003 | 26| 88 | 1 | 1 | 0 |
|s0004 | 45| 27 | 22 | 1 | 0 |
|s0004 | 10| 80 | 67 | 2 | 0 |
|s0004 | 71| 48 | 21 | 3 | 0 |
|s0004 | 58| 9 | 60 | 4 | 1 |
Hr_count: Represents the hour each subject stayed in the experiment
Label: This is my target variable when building my classifier. It Represents the flag that the subject received after staying in the experiment
I trained the data on an LSTM RNN model which was defined as below:
model = Sequential()
model.add(LSTM(100, activation='tanh', return_sequences=True, input_shape=(1, 25)))
model.add(LSTM(49, activation='tanh'))
model.add(Dense(1, activation='sigmoid'))
model.fit(
x_train, y_train,
validation_data=(x_test, y_test),
batch_size=32,
epochs=200)
Question:
Due to the sequential nature of the data, I would like to define a dynamic batch_size parameter to be the max number of Hr_count per subject in the training when fitting the model so that the LSTM can pick up the relationships between the data for each subject separately (Each batch will contain the data for each subject only). This will mean each batch contains samples for 1 subject, ordered by Hr_count.
The flexibility of having a dynamic batch_size does not seem to be available in Keras or TensorFlow v2.x (contrary to TensorFlow v1.x)...
How can I define the batch size to be dynamic for the batch_size parameter?