Neural network not learning after 1000 epochs to solve XOR problem

Viewed 177

I'm learning TensorFlow and I'm trying to solve the XOR problem. I created a 3 layers neural network to do that but after 500 or 1000 epochs its not learning at all. What am I doing wrong?

I'm using TensorFlow 2.3.0 in colab.research.google.

from tensorflow.keras.layers import Dense
from tensorflow.keras.losses import MeanSquaredError
from tensorflow.keras.optimizers import SGD
from tensorflow.keras.metrics import Accuracy
from tensorflow.keras import Sequential

import numpy as np



x = np.array([[0., 0.],
              [1., 1.],
              [1., 0.],
              [0., 1.]], dtype=np.float32)

y = np.array([[0.], 
              [0.], 
              [1.], 
              [1.]], dtype=np.float32)



model = Sequential()
model.add(Dense(2, activation='sigmoid'))
model.add(Dense(2, activation='sigmoid'))
model.add(Dense(1, activation='sigmoid'))
model.compile(optimizer='SGD', loss='mean_squared_error', metrics='accuracy')
model.fit(x, y, batch_size=1, epochs=1000, verbose=False)

pred = model.predict_on_batch(x)
print(pred)
1 Answers

Since you have mentioned hidden layer unit as 2 i.e Dense(2) which is not sufficient for the model to learn given the input of array having 2 inputs. I have included 16 Units, you can try experimenting with 32,64, etc units.

It's ideal to use the activation function ReLu for the hidden layer in the neural network. (Refer to Mark's comment for more details on this).
But for this use case, you can go without mentioning any activation function, but it takes more epochs to converge to a solution.

Below is the modified code, which predicts the correct output with a lesser number of epochs.

import tensorflow as tf
from tensorflow.keras.layers import Dense
from tensorflow.keras.losses import MeanSquaredError
from tensorflow.keras.optimizers import SGD
from tensorflow.keras.metrics import Accuracy
from tensorflow.keras import Sequential

import numpy as np



x = np.array([[0., 0.],
              [1., 1.],
              [1., 0.],
              [0., 1.]], dtype=np.float32)

y = np.array([[0.], 
              [0.], 
              [1.], 
              [1.]], dtype=np.float32)



model = Sequential()
model.add(Dense(16, activation='relu'))
model.add(Dense(1, activation='sigmoid'))
model.compile(optimizer='SGD', loss='mean_squared_error', metrics=['accuracy'])
model.fit(x, y, batch_size=1, epochs=500, verbose=False)

pred = model.predict(x).round()
print(pred) 

Output:

[[0.]
 [0.]
 [1.]
 [1.]]
Related