The network I am trying to build, using the tensorflow python library, looks like this:
A simple network with two nodes on the input layer, one on the output layer and two nodes on the hidden layer. Nodes z1 and z2 use the sigmoid activation function, and node y is just a linear combination of z1 and z2, without an activation function. The loss function is the mean squared error.
The function I am trying to learn is:
y = 0.26(x1^2 + x2^2) - 0.48*x1*x2 where xi ∈ [-10,10] for i = 1,2
Although it is non-linear, it's not a particularly complicated function, so it should be completely fine for this network to learn, given a training set of 100 samples generated using latin hypercube sampling on the domain - and in fact, I can do it with 3 lines of code in MATLAB:
rng(0, 'twister');
x_train = (repmat(-10, 100, 1) + repmat(10, 100, 1)...
.* lhsdesign(100, 2))';
y_train = 0.26 * (x_train(1, :).^2 + x_train(2, :).^2)...
- 0.48 .* x_train(1, :) .* x_train(2, :);
x_test = (repmat(-10, 10, 1) + repmat(10, 10, 1)...
.* lhsdesign(10, 2))';
y_test = 0.26 * (x_test(1, :).^2 + x_test(2, :).^2)...
- 0.48 .* x_test(1,:) .* x_test(2,:);
net = feedforwardnet(2, 'trainlm');
trained_net = train(net, x_train, y_train);
y_pred = trained_net(x_test);
perform(trained_net,y_test,y_pred) % output 0.0222 with rng seed 0
Which gets me pretty accurate predictions (MSE of 0.0222 with 100 training samples and 10 test samples), and even more accurate if the number of samples is increased.
I am having trouble recreating these results using the tensorflow python library though. I have looked at a lot of examples online and read through the documentation for the library, but most of the examples that I can find are for building classifier networks, which is not what I am trying to do, and when I try to adapt the examples to my purpose, it doesn't work.
The code I have written so far is:
import numpy as np
from tensorflow import keras
from tensorflow.keras import layers
import pyDOE as doe
np.random.seed(0)
def obj_fn(X):
return np.array([0.26 * (X[:, 0] ** 2 + X[:, 1] ** 2)
- 0.48 * X[:, 0] * X[:, 1]]).T
def MSE(y_hat, y):
return np.square(y_hat - y).mean()
x_train = -10 * np.ones((100, 2))\
+ 20 * np.ones((100, 2)) * doe.lhs(2, samples=100)
y_train = obj_fn(x_train)
x_test = -10 * np.ones((10, 2))\
+ 20 * np.ones((10, 2)) * doe.lhs(2, samples=10)
y_test = obj_fn(x_test)
model = keras.Sequential([
keras.Input(shape=(2), name='input'),
layers.Dense(2, activation='sigmoid', name='hidden'),
layers.Dense(1, name='output'),
])
model.compile(
optimizer='adam', # Optimizer
loss='mse',
)
history = model.fit(
x_train,
y_train,
epochs=1000)
y_pred = model.predict(x_test)
MSE(y_pred,y_test) # output average around 70-110 with rng seed 0
But when trained on 100 samples for 1000 epochs, and then tested on the test set of 10 samples, it only gets a mean squared error of 84.78, which is pretty terrible. I can get much better results if I increase the size of the hidden layer, but this should be quite a simple problem for this network to solve (and indeed MATLAB did it with no issues), so I am clearly doing something wrong somewhere. Can someone please help me work out what it is?
