I have this dataset containing training data and testing data , and I have to plot training and testing errors for the gradient descent algorithm with square and logistic losses. I'm a beginner, I'm a bit lost on what to do.
So far, I think I have managed to implement the gradient descent algorithm successfully, here is the one i did for the square loss:
import numpy as np
import scipy.io
import matplotlib.pyplot as plt
data = scipy.io.loadmat('data_orsay_2017.mat')
x0,x1=data['Xtrain'],data['Xtest']
y0,y1=data['ytrain'],data['Ytrain']
def gradientdescent(x,y,n,alpha,max_iterations): # n is the sample size, alpha is the learning rate
d = x.shape[1] # dimension of the data
theta = np.random.random(d)
error = []
for j in range(max_iterations):
prediction = x.dot(theta)
cost = 1/(2*n)*sum((y[i,0]-prediction[i])**2 for i in range(n))
error.append(cost)
grad = (1/n) * sum((prediction[i] - y[i,0])*x[i] for i in range(n))
theta-=alpha*grad
return (theta,error)
Now, i could simply run the algorithm on x0,y0 and x1,y1 and plot the errors, but I don't think that's what I'm supposed to do is it ? I assume I am supposed to treat the training data and the testing data differently, but I don't know how.