Pytorch DL model, updates normally with converging losses during training, but SAME OUTPUT values for all data (regression)

Viewed 45

I've only used Neural Networks in CNNs and RNNs, but this is my first time using it as a regression task.

There are 30000 sets of data. Each data has 50 input features, and I must predict 14 output features for each.

So, my goal is to make a prediction for about 30000 datasets, so that would make my task : IN - 30000 data X 50 features -> OUT - 30000 predictions X 14 features

these are my hyperparameters :

input_size = 50
hidden_size = 40
num_epochs = 7
learning_rate =1.00E-03
output_size=15
batch_size=30

My code worked fine, losses converged at every iteration/epoch.

However, for some reason I noticed my output which was expected of 30000 prediction (rows) X 14 feature (columns) returned just a same 1 tensor X 14 feature X 30000 times.

Like this :

[[  1.3311,   1.0411,   0.9971,  13.6349,  31.4082,  16.5008,   3.2034,
         -26.2985, -26.3108, -22.4322,  24.3007, -26.2376, -26.2337, -26.2369],
        [  1.3311,   1.0411,   0.9971,  13.6349,  31.4082,  16.5008,   3.2034,
         -26.2985, -26.3108, -22.4322,  24.3007, -26.2376, -26.2337, -26.2369],
        [  1.3311,   1.0411,   0.9971,  13.6349,  31.4082,  16.5008,   3.2034,
         -26.2985, -26.3108, -22.4322,  24.3007, -26.2376, -26.2337, -26.2369]]

Just imagine this outcome, only with 3000 rows. (I don't even understand how I reached converging loss) example of same output values for all data

I tried to track where this problem started, and it seems it's been happening during training as well.


    # 5. Training loop
    n_total_steps = len(DS)
    n_iterations = -(-n_total_steps // batch_size) # ceiling division
    training_loss=[]

    loss_fn = nn.MSELoss()

    trainloader = torch.utils.data.DataLoader(
                          DS, 
                          batch_size=batch_size, shuffle = True) 
    testloader = torch.utils.data.DataLoader(
                          TS,
                          batch_size=batch_size)
      
    for epoch in range(num_epochs):
        print('\n')

        for i, (data, target) in enumerate(trainloader): 
            data, target = data.to(device), target.to(device)
            outputs = model(data)
            loss = torch.sqrt(loss_fn(outputs, target))
            training_loss.append(loss.item())

        # 5.5 Backward pass
            opt.zero_grad() # 5.6 Empty the values in the gradient attribute, or model.zero_grad()
            loss.backward() # 5.7 Backprop
            opt.step() # 5.8 Update params

        # 5.9 Print loss
            if (i+1) % 100 == 0:
                print(f'Epoch {epoch+1}/{num_epochs}, Iteration {i+1}/{n_iterations}, Loss={loss.item():.4f} ')

    Epoch 1/7, Iteration 100/1321, Loss=1.5157 
    tensor([[  1.4186,   1.1157,   1.0471,  13.5818,  31.3844,  16.5334,   3.1015,
             -26.3141, -26.2974, -22.4117,  24.3678, -26.2477, -26.2577, -26.2387],
            [  1.4186,   1.1157,   1.0471,  13.5818,  31.3844,  16.5334,   3.1015,
             -26.3141, -26.2974, -22.4117,  24.3678, -26.2477, -26.2577, -26.2387],
            [  1.4186,   1.1157,   1.0471,  13.5818,  31.3844,  16.5334,   3.1015,
             -26.3141, -26.2974, -22.4117,  24.3678, -26.2477, -26.2577, -26.2387],
            [  1.4186,   1.1157,   1.0471,  13.5818,  31.3844,  16.5334,   3.1015,
             -26.3141, -26.2974, -22.4117,  24.3678, -26.2477, -26.2577, -26.2387],
            [  1.4186,   1.1157,   1.0471,  13.5818,  31.3844,  16.5334,   3.1015,
             -26.3141, -26.2974, -22.4117,  24.3678, -26.2477, -26.2577, -26.2387],
             ....
Epoch 1/7, Iteration 300/1321, Loss=0.9697 
tensor([[  1.3142,   1.0427,   0.9661,  13.6267,  31.2973,  16.5265,   3.1028,
             -26.2207, -26.2468, -22.3516,  24.4410, -26.1698, -26.1708, -26.1715],
            [  1.3142,   1.0427,   0.9661,  13.6267,  31.2973,  16.5265,   3.1028,
             -26.2207, -26.2468, -22.3516,  24.4410, -26.1698, -26.1708, -26.1715],
            [  1.3142,   1.0427,   0.9661,  13.6267,  31.2973,  16.5265,   3.1028,
             -26.2207, -26.2468, -22.3516,  24.4410, -26.1698, -26.1708, -26.1715],
            [  1.3142,   1.0427,   0.9661,  13.6267,  31.2973,  16.5265,   3.1028,
             -26.2207, -26.2468, -22.3516,  24.4410, -26.1698, -26.1708, -26.1715],
            [  1.3142,   1.0427,   0.9661,  13.6267,  31.2973,  16.5265,   3.1028,
             -26.2207, -26.2468, -22.3516,  24.4410, -26.1698, -26.1708, -26.1715],
            [  1.3142,   1.0427,   0.9661,  13.6267,  31.2973,  16.5265,   3.1028,
             -26.2207, -26.2468, -22.3516,  24.4410, -26.1698, -26.1708, -26.1715],

To elaborate : my prediction model's output returns same value for every different 3000 input data. It updates them altogether as well, not updating each rows separately!

I don't understand how it is reaching low loss and convergence.

NN code :

class MyDataset(Dataset) :

      def __init__(self, file_name) :
          train_df = pd.read_csv(file_name)
          x = train_df.filter(regex='X') # Input : X Featrue
          y = train_df.filter(regex='Y') # Output : Y Feature
          self.train_x = torch.tensor(x.values,dtype=torch.float32)
          self.train_y = torch.tensor(y.values,dtype=torch.float32)
  
      def __len__(self) :
          return len(self.train_y)
  
      def __getitem__ (self,idx) :
          return self.train_x[idx],self.train_y[idx]
    
class LGNN(nn.Module):
      def __init__(self, input_size, hidden_size, output_size):
          super().__init__()
          self.layer1 = nn.Linear(input_size, hidden_size)
          self.relu = nn.Tanh()
          self.layer2 = nn.Linear(hidden_size, hidden_size)
          self.layer3 = nn.Linear(hidden_size, hidden_size)
          self.layer4 = nn.Linear(hidden_size, output_size)

      def forward(self, x):
          out = self.layer1(x)
          out = self.relu(out)
          out = self.layer2(out)
          out = self.relu(out)
          out = self.layer3(out)
          out = self.relu(out)
          out = self.layer4(out)
          return out

      # 4.1 Create NN model instance
      model = LGNN(input_size, hidden_size, output_size).to(device) #to(device)는 GPU
      model.apply(reset_weights)
      # 4.2 Loss and Optimiser
      opt = optim.Adam(model.parameters(), lr=learning_rate)
      loss_fn = nn.MSELoss()
  • Also, the model has gone through k-fold and validation. No overifitting or other problems found :(
1 Answers
Related