Why does my class-weighted model perform so strangely/bad?

Viewed 37

Here are the accuracy and loss plots for the class-weighted version:
enter image description hereenter image description here

Here are the accuracy and loss plots for the unweighted version:
enter image description hereenter image description here

Here is the code. The only difference in the above two versions is that one calls the class weights dictionary and one doesn't. (General advice about how this is set up is also welcome -- as you can see I am very new to this!)

from tensorflow import keras
from keras import optimizers
from keras.applications.resnet_v2 import ResNet50V2
from tensorflow.keras.preprocessing import image
from tensorflow.keras.models import Model
from tensorflow.keras import layers
from tensorflow.keras.layers import Dense, GlobalAveragePooling2D, Rescaling, Conv2D, MaxPool2D, Flatten


#Create datasets
train_ds = tf.keras.preprocessing.image_dataset_from_directory(
    '/content/drive/MyDrive/Colab Notebooks/train/All classes/',
    labels="inferred",
    label_mode="int",
    validation_split=0.2,
    seed=1337,
    subset="training",
)

val_ds = tf.keras.preprocessing.image_dataset_from_directory(
    '/content/drive/MyDrive/Colab Notebooks/train/All classes/',
    labels="inferred",
    label_mode="int",
    validation_split=0.2,
    seed=1337,
    subset="validation",
)

test_ds = tf.keras.preprocessing.image_dataset_from_directory(
    '/content/drive/MyDrive/Colab Notebooks/test/All classes/',
    labels="inferred",
    label_mode="int",
)

#Import ResNet
base_model = ResNet50V2(weights='imagenet', include_top=False)

#Create basic network to append to ResNet above
x = base_model.output
x = Rescaling(1.0 / 255)(x)
x = Conv2D(32, kernel_size=(3, 3), activation='relu', input_shape=(256,256,3), padding="same")(x)
x = MaxPool2D(pool_size=(2, 2), strides=2)(x)
x = Conv2D(64, kernel_size=(3, 3), activation='relu')(x)
x = MaxPool2D(pool_size=(2, 2), strides=2)(x)
x = GlobalAveragePooling2D()(x)
predictions = Dense(units=5, activation='softmax')(x)

#Merge the models
model = Model(inputs=base_model.input, outputs=predictions)

#Freeze ResNet layers
for layer in base_model.layers:
    layer.trainable = False

#Compile
model.compile(optimizer=keras.optimizers.Adam(1e-3), loss='sparse_categorical_crossentropy', metrics=['accuracy'])

#These are the weights. They are derived here from their numbers in the train dataset -- there are 25,811 
#files in class 0, 2444 files in class 1, etc. This dictionary was not called for the unweighted version.
class_weight = {0: 1.0,
                1: 25811.0/2444.0,
                2: 25811.0/5293.0,
                3: 25811.0/874.0,
                4: 25811.0/709.0}

#Training the model
model_checkpoint_callback = tf.keras.callbacks.ModelCheckpoint(
    filepath='/content/drive/MyDrive/Colab Notebooks/ResNet/',
    save_weights_only=False,
    mode='auto',
    save_best_only=True,
    save_freq= 'epoch')

history = model.fit(
          x=train_ds,
          epochs=30,
          class_weight=class_weight,
          validation_data=val_ds,
          callbacks=[model_checkpoint_callback]
)

#Evaluating
loss, acc = model.evaluate(test_ds)
print("Accuracy", acc)

Also, some other questions:

  1. Should the metrics=['accuracy'] actually be metrics=['sparse_categorical_accuracy']?
  2. Should class_weight=class_weight actually be sample_weight=sample_weight? I couldn't tell the difference in the documentation, although most examples seem to use class_weight.
  3. I only used padding in one Conv2D layer, and this was a bodge to force the whole thing to actually compile. Should I have been more consistent and used it for the other one too?
  4. On that note, are there other ways my simple appended CNN model (it was called 'predictions') could be laid out to make better sense?
  5. Ah yes, before I forget -- as you can see from the above code, I didn't preprocess the data in accordance with the keras guidance for ResNet. I figured it probably wouldn't make that much of a difference (but also because I was having trouble trying to implement it). Would that be worth looking into? I suppose the unweighted model shows a very high accuracy... probably too high now that I'm looking at it... oh, dear.

I shall be so very thankful for any advice!

0 Answers
Related