How can I estimate value of weights (theta) in Logistic Regression?

Viewed 278

I am training a dataset in which I have to predict the status of loan (yes/no) based on features given like gender, dependents, total income, loan amount, loan duration, graduation, married, property, self-employed. I have written a logistic regression model for it. The problem I am facing is that the label (y_label), which the probability of getting a loan lies between [0,1], returns equal to 1 (or 0.9999) for all entries. The sigmoid function has been used to evaluate the probability of the dot product of the weights (theta) and feature vector(NumPy array of features of each individual entry). Can you please tell me why the final probability is coming out to be equal to 1 when theta (vector) has a large initial value and similarly it is equal to 0 when theta has small initial value??

My code looks like this It is a sigmoid function to calculate the probability of y predicted labels

def calc_sigmoid(z):
    p=1/(1+ np.exp(-z))
    p=np.minimum(p, 0.9999)
    p = np.maximum(p, 0.0001)
    #print("value of sigmoid", p)
    return p

Function to calculate the cost:

def calc_cost_func(theta,x):
    y=np.dot(theta,np.transpose(x))
    return calc_sigmoid(y)

Function to calculate error:

def calc_error(y_pred, y_label):
    len_label=len(y_label)
    cost= (-y_label*np.log(y_pred) - (1-y_label)*np.log(1-y_pred)).sum()/len_label
    return cost

Function to calculate gradient descent:

def gradient_descent(y_pred,y_label,x, learning_rate, theta):
    len_label=len(y_label)
    J= (-(np.dot(np.transpose(x),(y_label-y_pred)))/len_label)
    theta-= learning_rate*J
    return theta

Function to train data:

def train(y_label,x, learning_rate, theta, iterations):
    list_cost=[]
    for i in range(iterations):
        y_pred=calc_cost_func(theta,x)
        
        theta=gradient_descent(y_pred,y_label,x, learning_rate, theta)
        if i%100==0:
            print("\n iteration",i)
            print("y_label:",y_pred)
            print("theta:",theta)
    
        cost=calc_error(y_pred, y_label)
        list_cost.append(cost)
    
    return theta, cost

Extracting data from data frame loan:

'''
    theta: array(1,no. of features)
    theta_0: array(1,length of data)
    x: array(no. of features, length of data)
    y_label, y_pred: array(length of data)
'''

x_label=loan.iloc[:500, 1:10].values

x_rows, x_columns= x_label.shape

z = np.ones((x_rows,1), dtype=float)

x_label=np.append(x_label,z,axis=1)

x_label=x_label.astype(float)

y_label=loan.loc[:499,"Status_New" ].values

y_label=y_label.astype(float)

#theta=np.array([0.0000005,0.0000000455547,0.000000222203,0.0000066005,0.000000022505,0.0000000025059,0.000000002585,0.000025500049,0.00000000034,0.00000000068])
theta=np.array([50.0,20.0,40.0,10.0,5.0,35.0,12.0,40.0,69.0,40.5])

train(y_label,x_label,0.2,theta,1000)

Please help me why am I getting the probability equal to 1 for large theta values and 0 for small theta values?

0 Answers
Related