MATLAB Backpropagation Algorithm not functioning as expected

Viewed 213

I am attempting to write a Multi-Layer Perceptron Network inside MATLAB to help me better understand the calculus required for backpropagation.

The aim is so provide the network with XOR data (where upper-right and lower-left quadrant data is class 1 and the remaining quadrants class 0), train the network on this data, and then test it on new data.

My problem is that my loss curve looks very very strange:

Loss curve

It appears to bounce between very low error very high error and converge in the middle to a pretty poor error.

I was wondering if someone could check that I have correctly implemented the chain rule in MATLAB syntax.

The MLP network is structured as follows: Input-layer has 2 neurons, 1 hidden-layer with 2 neurons, and 1 output neuron.

Here is the MATLAB code:

%Create XOR Dataset

x1pos = rand(500,1);
x1neg = -rand(500,1);
x1 = [x1pos; x1neg];
p = randperm(length(x1));
x1 = x1(p);
x2pos = rand(500,1);
x2neg = -rand(500,1);
x2 = [x1pos; x1neg];
p = randperm(length(x2));
x2 = x2(p);

Data = [x1 x2];
TrainingData = Data(1:800,:);
TestData = Data(801:length(Data),:);

T = gt((Data(:,1).*Data(:,2)),0); %Create class label for data and assign to matrix T

%Neural Net

%Training

W1 = rand(2,2); %Initialize random weights
W2 = rand(1,2); %Initialize random weights

B1 = rand(2,1); %Initialize random biases
B2 = rand(1,1); %Initialize random biases

n = 0.05; %Set Learning Rate


    for i = 1:800
        %Fwd Pass
    
        x1 = Data(i,1);
        x2 = Data(i,2);
        X = [x1; x2];
        A1 = W1*X + B1;
        H1 = sigmoid(A1);
        A2 = W2*H1 + B2;
        Y = sigmoid(A2);
    
        %Loss
        
        Loss = (Y-T(i))*(Y-T(i));
    
        scatter(i, Loss)
        hold on;
    
        %Backpropagation
    
        dEdY = 2*(Y-T(i)); %The partial derivative of the loss with respect to the output
        dYdA2 = Y*(1-Y); %The partial derivative of the output with respect to the hidden layer output
        dA2dH1 = W2.'; %The partial derivative of the hidden layer output with respect to the first layer activations
        dH1dA1 = H1.*(1-H1); %The partial derivative of the first layer activations with respect to the first layer output
        
        %Chain Rule
        
        dEdW2 = dEdY.*dYdA2.*W2.';
        dEdW1 = dEdY.*dYdA2.*dA2dH1.*dH1dA1.*W1.';
        dEdB2 = dEdY.*dYdA2;
        dEdB1 = dEdY.*dYdA2.*dA2dH1.*dH1dA1;
    
        %Update Weights
    
        W2 = (W2.' - n.*dEdW2).';
        W1 = (W1.' - n.*dEdW1).';
    
        %Update Biases
    
        B2 = B2 - n.*dEdB2;
        B1 = B1 - n.*dEdB1;
    
        %Next training loop
    end

%Testing

for i = 801:1000
    x1 = Data(i,1);
    x2 = Data(i,2);
    X = [x1; x2];
    A1 = W1*X + B1;
    H1 = sigmoid(A1);
    A2 = W2*H1 + B2;
    Y = sigmoid(A2);
end

function o = sigmoid(input)
    o = [];
    for i = 1:length(input)
        o = [o; 1/(1+exp(-input(i)))];
    end
end
0 Answers
Related