I am working on a problem of time-series segmentation of sleep sensor data by using U-Net (CNN-based) in a tensorflow(keras) implementation and I am getting some really strange behavior. For a 10-fold cross-validation run, typically 9 folds work as expected, with one or two folds containing 0 positive predictions. I am at a complete loss at this point as to why some of the folds refuse to learn anything and seem to get stuck in a local minima after 1 epoch... I am really looking for ANY input that could point me in the right direction, thanks!
For example I receive:
TP=87153
TN=1889185
FP=35217
FN=93021
for a typical fold, but in a problematic one:
TP=0
TN=1912697
FP=0
FN=191879
As you can see, the model predicts no positives for these runs.
Stats for 10 folds; recall/precision is 0.000 for the problematic fold #2 (technically undefined):
| fold | acc | precision | recall | f1 | specificity |
|---|---|---|---|---|---|
| 1 | 0.895 | 0.752 | 0.629 | 0.685 | 0.954 |
| 2 | 0.862 | 0.000 | 0.000 | 0.000 | 1.000 |
| 3 | 0.903 | 0.727 | 0.592 | 0.653 | 0.960 |
| 4 | 0.894 | 0.760 | 0.531 | 0.625 | 0.966 |
| 5 | 0.893 | 0.805 | 0.561 | 0.661 | 0.969 |
| 6 | 0.901 | 0.755 | 0.583 | 0.658 | 0.963 |
| 7 | 0.900 | 0.760 | 0.518 | 0.616 | 0.970 |
| 8 | 0.865 | 0.857 | 0.522 | 0.649 | 0.973 |
| 9 | 0.891 | 0.779 | 0.572 | 0.659 | 0.963 |
| 10 | 0.895 | 0.806 | 0.592 | 0.683 | 0.967 |
I have tried these measures with no effect:
- Reducing the size of the model (# of nodes/feature channels)
- Reducing the number of convolutional layers in U-Net
- Balancing the dataset (It is quite unbalanced)
- Removing dropout layers
- Upgrading to tensorflow 2.3.
- Running 10 runs with different randomSeeds, this simply shifts the problem to another fold, and the zero-fold behavior appeared in 7/10 of the runs. This tells me that the error is not data-dependent.
- Reducing the size of the data-set to give me some clues. This makes the problem worse, and introduces more problematic zero-positive folds (up to 5/10), but the problematic folds remain a superset of the problematic folds in the runs with more data, which is further evidence that it is not data-dependent.
A problematic fold's run will have this behavior where the loss starts extremely high and gradually goes down, but val_accuracy and accuracy stay the same (except for the first epoch). I am using ReLu as the activation function, which I think should prevent vanishing gradients(?):
Epoch 1/50
397/397 [==============================] - 13s 33ms/step - loss: 0.6851 - accuracy: 0.9010 - val_loss: 0.6777 - val_accuracy: 0.8963
Epoch 2/50
397/397 [==============================] - 12s 30ms/step - loss: 0.6688 - accuracy: 0.9207 - val_loss: 0.6628 - val_accuracy: 0.8963
Epoch 3/50
397/397 [==============================] - 12s 30ms/step - loss: 0.6532 - accuracy: 0.9207 - val_loss: 0.6484 - val_accuracy: 0.8963
Epoch 4/50
397/397 [==============================] - 12s 30ms/step - loss: 0.6381 - accuracy: 0.9207 - val_loss: 0.6345 - val_accuracy: 0.8963
Epoch 5/50
397/397 [==============================] - 12s 30ms/step - loss: 0.6235 - accuracy: 0.9207 - val_loss: 0.6210 - val_accuracy: 0.8963
A normal fold's epochs by contrast:
Epoch 1/50
397/397 [==============================] - 12s 31ms/step - loss: 0.2710 - accuracy: 0.9195 - val_loss: 0.2063 - val_accuracy: 0.8963
Epoch 2/50
397/397 [==============================] - 12s 31ms/step - loss: 0.1671 - accuracy: 0.9195 - val_loss: 0.1953 - val_accuracy: 0.8963
Epoch 3/50
397/397 [==============================] - 12s 31ms/step - loss: 0.1613 - accuracy: 0.9342 - val_loss: 0.1915 - val_accuracy: 0.9302
Epoch 4/50
397/397 [==============================] - 12s 31ms/step - loss: 0.1582 - accuracy: 0.9436 - val_loss: 0.1867 - val_accuracy: 0.9318
Epoch 5/50
397/397 [==============================] - 12s 31ms/step - loss: 0.1567 - accuracy: 0.9444 - val_loss: 0.1871 - val_accuracy: 0.9316
Hyperparameters:
validationFraction: 0.33
batchSize: 500
numFolds: 10
numEpochs: 50
I would really appreciate any thoughts or anecdotal insights if you have encountered something similar.