naive bayes algorithm stuck in accuracy of 50%

Viewed 79

I am having trouble writing a naive bayes algorithm from scratch. I am going for Gaussian Density, at it seems to have a higher accuracy than simple Naive Bayes, but I am still only getting 50% accuracy (i.e., the program isn't actually predicting anything...it's essentially guessing, as there is a 50 percent chance of guessing class 1 or 2).

A brief explanation of my data: I am using the IMDB movie review dataset, with the sentiment column labelled as 1 for positive, 0 for negative. I have transformed x and x_test into a tfidf vector (using fit_transform for x, and transform for x_test) and a given row in x/x_test might look as follows: [0.0, 0.0, 0.07955215256290477, 0.0, 0.0, 0.0, 0.0, 0.050875508532914046, 0.0, 0.0, 0.0] (but much longer obviously)

I have tried and tried, and am at a loss on where my math is wrong - everything I have tried has gotten me 55% accuracy or less. The desired accuracy is 80%+, as an average naive bayes model. I believe the issue is in regards to the math part of the algorithm; any feedback is appreciated.

  means = []
  std = []
  means = [0 for i in range(len(x[0]))] 
  std = [0 for i in range(len(x[0]))] 


  for j in range(len(x[0])):
    means[j] = (x[j].sum())/float(len(x))
    vars[j] = x[j].std()

  return means, std
def gaussian_density(x, mean, std):     
  expo = np.exp(-((x-mean)**2 / (2*std**2)))
  return (1 / (np.sqrt(2 * np.pi) * std)) * expo

def predict(x, y, x_test, y_test):

  pos, neg = group_by_class(x, y)

  num_pos = len(pos)
  num_neg = len(neg)
  result = []

  sums_pos = get_sums(pos)
  sums_neg = get_sums(neg)

  # mean, var
  mean_pos, std_pos = calculations(pos)
  mean_neg, std_neg = calculations(neg)

  mean, std = calculations(x)


  index = 0
  number_correct = 0

  for i in range(0, len(x_test)):
    prior_pos = np.log(num_pos / len(x))
    prior_neg = np.log(num_neg / len(x))
    prob_pos = 1
    prob_neg =1

    wc = x_test[i]

    for word_index in range(0, len(wc)): 

      prob_pos += (gaussian_density(sums_pos[word_index], mean_pos[word_index], std_pos[word_index]))
      prob_neg += (gaussian_density(sums_neg[word_index], mean_neg[word_index], std_neg[word_index]))
prob_pos *=prior_pos
prob_neg *= prior_neg


    if prob_pos >= prob_neg:
      result.append(1)

      # For my own checks so I dont have to wait for the entire thing to finish
      if y_test.iloc[i] == 1:
        number_correct+=1

        print(number_correct/(index+1))
    else:
      result.append(0)

      # for my own checks
      if y_test.iloc[i] == 0:
        number_correct+=1
        print(number_correct/(index+1))

    
  
    index += 1

  return result
0 Answers
Related