How to predict one value based on a series input of 2D shape correctly?

Viewed 296

I am using an encoder-decoder architecture, with 3 layers each in the encoder and decoder and 128 neurons in each hidden layer. The inputs are in a 2D form: the first column has the days and the second column has the time series dependent on the days (shape:(5780, 100, 2)). The output is a single value among the values of the first column, that represents a particular day when the breaking point occurs (shape:(5780, 1, 1)). The breaking point is one of the time dependent values i.e. the second column.

A better picture of the input is:

array([[  0.        ,   1.        ],
       [  2.        ,   1.14469799],
       [  4.        ,   1.35245666],
       ...,
       [ 96.        ,   1.80030942],
       [ 98.        ,   1.79964733],
       [100.        ,   1.9898739]])

Where the days are in the first column and the corresponding points measured are in the second column.

The output is just a single value, which represents the day at which the breaking point occurs:

array([[1108.]])

The problem is that after training, the output on all different test data is almost exactly the same, i.e. it gives the same day for breaking points of all different materials (with just negligible changes in the decimal places). I have experimented with high and low learning rates (ranging 1e-2 to 1e-5), the number of training epochs (300 to 3000). I have also varied the number of layers and the neurons per layer.

What I haven't done is batch normalization or any kind of normalization, but I had done some operations with the same data having the same gradients, and it worked perfectly fine.

The architecture that I'm using here is as follows:

nodes = 128
drp = 0.01

# Defining input layers and shapes
input_train = Input(shape = (complete_inputs.shape[1], complete_inputs.shape[2]))
output_train = Input(shape= (kp_targets.shape[1], kp_targets.shape[2]))

# Masking layer
masking_layer = Masking(mask_value=0, input_shape = input_train.shape)(input_train)

# Encoder layer. For simple S2S model, we only need the last state_h and the last state_c.
enc_first_layer = Bidirectional(LSTM(nodes, dropout=drp, return_sequences=True, return_state=True))(masking_layer)
enc_first_layer, enc_fwd_h1, enc_fwd_c1, enc_back_h1, enc_back_c1 = Bidirectional(LSTM(nodes, dropout=drp, return_sequences=True, return_state=True))(enc_first_layer)
enc_stack_h, enc_fwd_h2, enc_fwd_c2, enc_back_h2, enc_back_c2 = Bidirectional(LSTM(nodes, dropout=drp, return_sequences=True, return_state=True))(enc_first_layer)

enc_last_h1 = concatenate([enc_fwd_h1, enc_back_h1])
enc_last_h2 = concatenate([enc_fwd_h2, enc_back_h2])
enc_last_c1 = concatenate([enc_fwd_c1, enc_back_c1])
enc_last_c2 = concatenate([enc_fwd_c2, enc_back_c2])


# RepeatVector layer (using only the last hidden state of encoder)
rv = RepeatVector(output_train.shape[1])(enc_last_h2)

# Stacked decoder layer for alignment score calculation (using the last hidden state of encoder)
dec_stack_h = Bidirectional(LSTM(nodes, dropout=drp, return_state=False, return_sequences=True))(rv, initial_state=[enc_fwd_h1, enc_fwd_c1, enc_back_h1, enc_back_c1])
dec_stack_h = Bidirectional(LSTM(nodes, dropout=drp, return_state=False, return_sequences=True))(dec_stack_h)
dec_stack_h = Bidirectional(LSTM(nodes, dropout=drp, return_state=False, return_sequences=True))(dec_stack_h, initial_state=[enc_fwd_h2, enc_fwd_c2, enc_back_h2, enc_back_c2])


# Attention layer (uses STACKED encoder output and dots it with stacked decoder output)
attention_ = dot([dec_stack_h, enc_stack_h], axes=[2,2])
attention_ = Activation('softmax')(attention_)

# Calculating the context vector
context = dot([attention_, enc_stack_h], axes=[2,1])

# Concat the context vector and stacked hidden states of decoder, and use it as input to the last dense layer
dec_combined_context = concatenate([context, dec_stack_h])


# Output Timedistributed dense layers
out = TimeDistributed(Dense(nodes/2, activation='relu'))(dec_combined_context)
out = TimeDistributed(Dense(output_train.shape[2], activation='linear'))(dec_combined_context)

# Compile model
model_attn = Model(inputs=input_train, outputs=out)
opt = optimizers.Adam(learning_rate=0.004)
model_attn.compile(optimizer=opt, loss=masked_mae)

What could be going wrong here?

Just to have a broader view on the problem, I also had the following question in mind: Is this model an overkill? Is there another machine/deep learning model that is more suited to predicting this kind of an output with the data that I have?

I have been at this problem for a week with no improvements, so any help will be greatly appreciated.

EDIT 1: Tried Normalization using StandardScaler and a simpler architecture. No improvements so far. Following is the structure with the commented-out sections implemented in all possible combinations..

nodes = 130 # Tried with 10/30/40/80

model_attn = Sequential()
#model_attn.add(Masking(mask_value=0, input_shape = (complete_inputs.shape[1], complete_inputs.shape[2])))

#model_attn.add(Bidirectional(LSTM(nodes, dropout=0.1, return_sequences=True)))
#model_attn.add(Bidirectional(LSTM(nodes, dropout=0.1, return_sequences=True)))
model_attn.add(Bidirectional(LSTM(nodes, dropout=0.1, return_sequences=False)))

model_attn.add(Dense(1))
model_attn.compile(optimizer=optimizers.Adam(0.001), loss = 'MAE')


No drop in losses over time:

model_attn.fit(complete_inputs, kp_targets, batch_size=350, epochs=300, shuffle=True, validation_split=0.1, callbacks=[callback])

Epoch 1/300
11/11 [==============================] - 18s 2s/step - loss: 0.7930 - val_loss: 0.3486
Epoch 2/300
11/11 [==============================] - 16s 1s/step - loss: 0.7544 - val_loss: 0.5152
Epoch 3/300
11/11 [==============================] - 16s 1s/step - loss: 0.7406 - val_loss: 0.4794
Epoch 4/300
11/11 [==============================] - 16s 1s/step - loss: 0.7385 - val_loss: 0.5361
Epoch 5/300
11/11 [==============================] - 16s 1s/step - loss: 0.7367 - val_loss: 0.4821
Epoch 6/300
11/11 [==============================] - 16s 1s/step - loss: 0.7350 - val_loss: 0.5518
Epoch 7/300
11/11 [==============================] - 18s 2s/step - loss: 0.7344 - val_loss: 0.5151
Epoch 8/300
11/11 [==============================] - 17s 2s/step - loss: 0.7339 - val_loss: 0.5646
Epoch 9/300
11/11 [==============================] - 16s 1s/step - loss: 0.7380 - val_loss: 0.5277
Epoch 10/300
11/11 [==============================] - 16s 1s/step - loss: 0.7382 - val_loss: 0.4879
Epoch 11/300
11/11 [==============================] - 16s 1s/step - loss: 0.7367 - val_loss: 0.5367
Epoch 12/300
11/11 [==============================] - 16s 1s/step - loss: 0.7382 - val_loss: 0.4910
Epoch 13/300
11/11 [==============================] - 16s 1s/step - loss: 0.7354 - val_loss: 0.5244
Epoch 14/300
11/11 [==============================] - 16s 1s/step - loss: 0.7386 - val_loss: 0.5043
Epoch 15/300
11/11 [==============================] - 16s 1s/step - loss: 0.7329 - val_loss: 0.5421
Epoch 16/300
11/11 [==============================] - 16s 1s/step - loss: 0.7376 - val_loss: 0.5023
Epoch 17/300
11/11 [==============================] - 16s 1s/step - loss: 0.7346 - val_loss: 0.4539
.....
.....

Epoch 27/300
11/11 [==============================] - 15s 1s/step - loss: 0.7388 - val_loss: 0.5649
Epoch 28/300
11/11 [==============================] - 16s 1s/step - loss: 0.7329 - val_loss: 0.6575
Epoch 29/300
11/11 [==============================] - 16s 1s/step - loss: 0.7400 - val_loss: 0.5123
Epoch 30/300
11/11 [==============================] - 16s 1s/step - loss: 0.7336 - val_loss: 0.4965
Epoch 31/300
11/11 [==============================] - 16s 1s/step - loss: 0.7328 - val_loss: 0.5069
Epoch 32/300
11/11 [==============================] - 17s 2s/step - loss: 0.7320 - val_loss: 0.5274
Epoch 33/300
11/11 [==============================] - 17s 2s/step - loss: 0.7302 - val_loss: 0.5968
Epoch 34/300
11/11 [==============================] - 16s 1s/step - loss: 0.7354 - val_loss: 0.6161
....
....
....
Epoch 184/300
11/11 [==============================] - 16s 1s/step - loss: 0.7088 - val_loss: 0.8242
Epoch 185/300
11/11 [==============================] - 16s 1s/step - loss: 0.7034 - val_loss: 0.7799
Epoch 186/300
11/11 [==============================] - 16s 1s/step - loss: 0.7098 - val_loss: 0.8179
Epoch 187/300
11/11 [==============================] - 16s 1s/step - loss: 0.7066 - val_loss: 0.7854
Epoch 188/300
11/11 [==============================] - 16s 1s/step - loss: 0.7142 - val_loss: 0.8340
Epoch 189/300
11/11 [==============================] - 16s 1s/step - loss: 0.7123 - val_loss: 0.7197

With no particular order of increase or decrease in both the losses. I have also stopped training at the epochs with the minimum losses, but with no improvements.

UPDATE 1: There was a mistake in the application of the StandardScalar. After fixing it, it did seem to output different predictions for the test dataset!

UPDATE 2: For these types of predictions, CNNs are also a great choice. However, comparison between the two architectures still need to be done. Will update my findings here!

UPDATE 3: For these types of predictions, CNNs are a much better choice than LSTMs as the data pertains to more of a classification problem. Although with more layers of LSTMs and tuning of the hyperparameters can work, my experiments showed that for similar results, CNNs perform AT LEAST 12 times faster than LSTMs, and also possibly with lower memory usage.

0 Answers
Related