I am currently doing sentiment analysis project. I fit the vectorizer with my train data in dataframe format. Then I transform the test data with the same vectorizer but it returns nothing for me. I did check the TfidfVectorizer.get_feature_names() and the desired transform word already exist inside the features. What is wrong with my vectorizer?
Code:
vectorizer = TfidfVectorizer(analyzer=lambda x: x)
x = vectorizer.fit_transform(data['clean_text'])
print(vectorizer.get_feature_names()[9427])
# output sad
print(vectorizer.transform(["sad"]))
# empty result
print(vectorizer.transform(["sad"]).toarray())
# return a whole 0 array
Sample data format (dataframe)
sentiment clean_text
0 0 [respond, go]
1 1 [sooo, sad]
2 1 [bulli]
3 1 [leav, alon]
4 1 [cry]