LSTM model for time series forecasting does not train proprely for some data

Viewed 109

CONTEXT

I have a dataframe of monthly historical prices of market indices like so (all data comes from Bloomberg):

             MSCI World   S&P 500   ...   HFRX Event Driven   Gold Spot  
1969-12-31          100     92.06   ...                 NaN         NaN
1970-01-30        94.25     85.02   ...                 NaN         NaN
       ...          ...       ...   ...                 ...         ...
2021-07-31      3141.35   4395.26   ...           20459.292      143.77  
2021-08-31       3006.6   4522.68   ...           20614.276      134.06

I want to predict the value of each index for the next month with an LSTM NN (each index has its specially trained NN).

So a new LSTM model is initialized and trained on each of these time series (which all have from 300 to 1200 samples). This (Pytorch) LSTM model is the following:

class LSTMRegressor(nn.Module):
    def __init__(self, input_size, hidden_size,sequence_size,num_layers,dropout):
        super(LSTMRegressor,self).__init__()

        self.input_size = input_size
        self.hidden_size = hidden_size
        self.sequence_size = sequence_size
        self.num_layers=num_layers
        self.droput = dropout
        
        self.lstm = nn.LSTM(
            input_size=self.input_size,
            hidden_size=self.hidden_size,
            num_layers=num_layers,
            batch_first=True,
            dropout=dropout)
        self.linear = nn.Linear(in_features=hidden_size, out_features=1)

    def forward(self, x):
        lstm_out, self.hidden = self.lstm(x)
        y_pred = self.linear(lstm_out[:,-1,:])
        return y_pred

Loss function and optimizer:

criterion = nn.MSELoss()
optimizer = torch.optim.Adam(model.parameters(),lr=learning_rate)

My parameters are the following:

  • input_size = 1
  • hidden_size=150
  • num_layers=2
  • dropout=0
  • batch_size = 16
  • learning_rate = 0.001

RESULTS

For most of the indexes, the training seems to work well as there is only about a mean error of 0.5% in the testing set (see an exemple in first graph below). However, for some of the indexes, the training does not work (about 100% error) (see an exemple in second graph below).

The graphs show training/validation loss and mape (mean average percentage error). The vertical red line is simply the best epoch calculated by an "early stopping algorithm".

Model that trained successfully (test == validation): Training pretty much successful

Model that trained unsuccessfully (test == validation): Training that went wrong

QUESTIONS

  1. Why do all LSTM models seem not to overfit (I've tested with ten of thousands of epochs)?
  2. Why do some LSTM models do not train proprely? (they are not those with the least data)
  3. Why do the models that do not train proprely have such smooth curves?

Thank you very much for your help!

0 Answers
Related