my train test splitting of data is not working. unigrams show same words for all category

Viewed 133

output of unigrams for both categoryWhen splitting the text-classifier data into test train and using chi2 to find unigrams it shows same unigrams for both of my categories and training is also not done.

Using chi2 to find unigrams. Tried using another models to train but same result. Expected results should be as follows: I have two categories to classify cricket and football. and in each of those unigrams only specific and distinct words should be allowed. but for each cricket and football unigram same set of words are coming.

X = df.Content_Parsed
y = df.Category_Code
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.1, 
random_state = 42)
ngram_range = (1,2)
min_df = 1
max_df = 50.
max_features = 300

tfidf = TfidfVectorizer(encoding='utf-8',
                    ngram_range=ngram_range,
                    stop_words=None,
                    lowercase=False,
                    max_df=max_df,
                    min_df=min_df,
                    max_features=max_features,
                    norm='l2',
                    sublinear_tf=True)

features_train = tfidf.fit_transform(X_train).toarray()
labels_train = y_train
print(features_train.shape)
labels_train.head()

features_test = tfidf.transform(X_test).toarray()
labels_test = y_test
print(features_test.shape)
from sklearn.feature_selection import chi2
import numpy as np

for Product, category_id in sorted(category_codes.items()):
    features_chi2 = chi2(features_train, labels_train == category_id)
    indices = np.argsort(features_chi2[0])
    feature_names = np.array(tfidf.get_feature_names())[indices]
    unigrams = [v for v in feature_names if len(v.split(' ')) == 1]
    bigrams = [v for v in feature_names if len(v.split(' ')) == 2]
    print("# '{}' category:".format(Product))
    print("  . Most correlated unigrams:\n. {}".format('\n. '.join(unigrams[-5:])))
    print("  . Most correlated bigrams:\n. {}".format('\n. '.join(bigrams[-2:])))
    print("")
0 Answers
Related