There are multiple issues with your code. I have tried adding separate sections to explain them. Please go through all of them and do try out the code examples I have shown below.
1. Passing the samples/batch channel as the input dimension
You are passing the batch channel as the input shape for the dense layer. That is incorrect. Instead what you need to do is to pass the shape of each sample that the model should expect, in this case (128,128). The model automatically adds a channel in front for the batches to flow through the computation graph as (None, 128, 128), as shown in model.summary() below
2. 2D inputs for Dense layer.
Each of your samples (in this case total of 5000 samples) is a 2D matrix of shape 128,128. A dense layer can not directly consume it without flattening it first. (or using a different layer to be more suited to work with 2D/3D inputs as discussed later).
from tensorflow.keras import Sequential
from tensorflow.keras.layers import Dense, Flatten
model = Sequential()
model.add(Flatten(input_shape=(128,128)))
model.add(Dense(12, activation='relu'))
model.add(Dense(8, activation='relu'))
model.add(Dense(1, activation='sigmoid'))
model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])
model.summary()
Model: "sequential_6"
_________________________________________________________________
Layer (type) Output Shape Param #
=================================================================
flatten_3 (Flatten) (None, 16384) 0
_________________________________________________________________
dense_5 (Dense) (None, 12) 196620
_________________________________________________________________
dense_6 (Dense) (None, 8) 104
_________________________________________________________________
dense_7 (Dense) (None, 1) 9
=================================================================
Total params: 196,733
Trainable params: 196,733
Non-trainable params: 0
_________________________________________________________________
3. Using a different architecture for your problem.
"Is my NN too simplistic for this type of problem?"
It's not about the complexity of the architecture, but more about the type of layers that can process a certain type of data. In this case, you have images with single channels (128,128), which is a 2D input. Usually, color images have R, G, B channels, which end up as (128,128,3) shaped inputs.
The general practice is to use CNN layers for this.
An example of that is shown below -
from tensorflow.keras import Sequential
from tensorflow.keras.layers import Dense, Flatten, Conv2D, MaxPooling2D, Reshape
model = Sequential()
model.add(Reshape((128,128,1), input_shape=(128,128)))
model.add(Conv2D(5, 5, activation='relu'))
model.add(MaxPooling2D((2,2)))
model.add(Conv2D(10, 5, activation='relu'))
model.add(MaxPooling2D((2,2)))
model.add(Conv2D(20, 5, activation='relu'))
model.add(MaxPooling2D((2,2)))
model.add(Conv2D(30, 5, activation='relu'))
model.add(MaxPooling2D((2,2)))
model.add(Flatten())
model.add(Dense(12, activation='relu'))
model.add(Dense(8, activation='relu'))
model.add(Dense(1, activation='sigmoid'))
model.compile(loss='binary_crossentropy', optimizer='adam', metrics=['accuracy'])
model.summary()
Model: "sequential_13"
_________________________________________________________________
Layer (type) Output Shape Param #
=================================================================
reshape_5 (Reshape) (None, 128, 128, 1) 0
_________________________________________________________________
conv2d_16 (Conv2D) (None, 124, 124, 5) 130
_________________________________________________________________
max_pooling2d_16 (MaxPooling (None, 62, 62, 5) 0
_________________________________________________________________
conv2d_17 (Conv2D) (None, 58, 58, 10) 1260
_________________________________________________________________
max_pooling2d_17 (MaxPooling (None, 29, 29, 10) 0
_________________________________________________________________
conv2d_18 (Conv2D) (None, 25, 25, 20) 5020
_________________________________________________________________
max_pooling2d_18 (MaxPooling (None, 12, 12, 20) 0
_________________________________________________________________
conv2d_19 (Conv2D) (None, 8, 8, 30) 15030
_________________________________________________________________
max_pooling2d_19 (MaxPooling (None, 4, 4, 30) 0
_________________________________________________________________
flatten_9 (Flatten) (None, 480) 0
_________________________________________________________________
dense_23 (Dense) (None, 12) 5772
_________________________________________________________________
dense_24 (Dense) (None, 8) 104
_________________________________________________________________
dense_25 (Dense) (None, 1) 9
=================================================================
Total params: 27,325
Trainable params: 27,325
Non-trainable params: 0
_________________________________________________________________
To understand what Conv2D layers and MaxPooling layers do, do check out my well-accepted answer on Data Science stack exchange, or check out this blog. But a general way to mentally visualize what it's doing is the following diagram.
