I am having trouble writing a naive bayes algorithm from scratch. I am going for Gaussian Density, at it seems to have a higher accuracy than simple Naive Bayes, but I am still only getting 50% accuracy (i.e., the program isn't actually predicting anything...it's essentially guessing, as there is a 50 percent chance of guessing class 1 or 2).
A brief explanation of my data: I am using the IMDB movie review dataset, with the sentiment column labelled as 1 for positive, 0 for negative. I have transformed x and x_test into a tfidf vector (using fit_transform for x, and transform for x_test) and a given row in x/x_test might look as follows: [0.0, 0.0, 0.07955215256290477, 0.0, 0.0, 0.0, 0.0, 0.050875508532914046, 0.0, 0.0, 0.0] (but much longer obviously)
I have tried and tried, and am at a loss on where my math is wrong - everything I have tried has gotten me 55% accuracy or less. The desired accuracy is 80%+, as an average naive bayes model. I believe the issue is in regards to the math part of the algorithm; any feedback is appreciated.
means = []
std = []
means = [0 for i in range(len(x[0]))]
std = [0 for i in range(len(x[0]))]
for j in range(len(x[0])):
means[j] = (x[j].sum())/float(len(x))
vars[j] = x[j].std()
return means, std
def gaussian_density(x, mean, std):
expo = np.exp(-((x-mean)**2 / (2*std**2)))
return (1 / (np.sqrt(2 * np.pi) * std)) * expo
def predict(x, y, x_test, y_test):
pos, neg = group_by_class(x, y)
num_pos = len(pos)
num_neg = len(neg)
result = []
sums_pos = get_sums(pos)
sums_neg = get_sums(neg)
# mean, var
mean_pos, std_pos = calculations(pos)
mean_neg, std_neg = calculations(neg)
mean, std = calculations(x)
index = 0
number_correct = 0
for i in range(0, len(x_test)):
prior_pos = np.log(num_pos / len(x))
prior_neg = np.log(num_neg / len(x))
prob_pos = 1
prob_neg =1
wc = x_test[i]
for word_index in range(0, len(wc)):
prob_pos += (gaussian_density(sums_pos[word_index], mean_pos[word_index], std_pos[word_index]))
prob_neg += (gaussian_density(sums_neg[word_index], mean_neg[word_index], std_neg[word_index]))
prob_pos *=prior_pos
prob_neg *= prior_neg
if prob_pos >= prob_neg:
result.append(1)
# For my own checks so I dont have to wait for the entire thing to finish
if y_test.iloc[i] == 1:
number_correct+=1
print(number_correct/(index+1))
else:
result.append(0)
# for my own checks
if y_test.iloc[i] == 0:
number_correct+=1
print(number_correct/(index+1))
index += 1
return result