I would like to have an image embedding to understand which images the network is seing as closer and which one seemed to be very different for him. First, I wanted to use Tensorboard callbacks in Keras, but the documentation are not clear enough for me, and I could not find any useful examples to reproduce it. Hence, to make sure to understand what I am doing, I preferred to make the embedding myself.
To do so, I planed to download the model already trained on my data, removed the last layers (last dropout and dense layer), and predict on the validation images to get the features associated to each images. Then I would simply do a PCA on these features and plot the images according to their first three principal component values.
But I think I misunderstood something, as when I remove the last layers the model predictions are still of the size of the number of classes, but to me it should be of the size of the last layer, which is 128 in my case.
Below is the code for clarification (where I just put the lines which seems useful to answer the question, but do not hesitate to ask for more details):
#model creation
base_model = applications.inception_v3.InceptionV3(include_top=False,
weights='imagenet',
pooling='avg',
input_shape=(img_rows, img_cols, img_channel))
#Adding custom Layers
add_model = Sequential()
add_model.add(Dense(128, activation='relu',input_shape=base_model.output_shape[1:],
kernel_regularizer=regularizers.l2(0.001)))
add_model.add(Dropout(0.60))
add_model.add(Dense(2, activation='sigmoid'))
# creating the final model
model = Model(inputs=base_model.input, outputs=add_model(base_model.output))
Then I trained the model on a dataset having two classes, and loade the model plus its weight to produce features:
model = load_model(os.path.join(ROOT_DIR,'model_1','model_cervigrams_all.h5'))
#remove the last two layers
#remove dense_2
model.layers[-1].pop()
#remove dropout_1
model.layers[-1].pop()
model.summary() # last alyer output shape is : (None, 128), so the removal worked
#predict
model.predict(np.reshape(image,[1,image.shape[0],image.shape[1],3])) #output only two values
Where am I wrong? Would you have any recommendations?