ValueError: Found input variables with inconsistent numbers of samples: [40000, 10000]

Viewed 34

I am getting the error in this line of code (X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)

Error: ValueError: Found input variables with inconsistent numbers of samples: [40000, 10000]. Seems like after the vectorization, the array size gets change and that does not match with y. Seeking support in resolving the error. Thanks in advance

Output: (10000, 4) (10000,) (40000, 1500) (10000,)

import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.feature_extraction.text import CountVectorizer

# Import dataset:
dataset = pd.read_excel(r"C:\Users\HPS1RT\Downloads\test\Safety_Prediction.xlsx", nrows=10000)
dataset[["Safety"]] *= 1

# Assign values to the X and y variables:
X = dataset.iloc[:, :-1].values
y = dataset.iloc[:, 4].values
print(X.shape)
print(y.shape)
    
#vectorization
vectorizer = CountVectorizer(max_features=1500, min_df=5, max_df=0.7)
X = vectorizer.fit_transform(X.ravel()).toarray()

from sklearn.feature_extraction.text import TfidfTransformer
tfidfconverter = TfidfTransformer()
X = tfidfconverter.fit_transform(X).toarray()
print(X.shape)
print(y.shape)

# Split dataset into random train and test subsets:
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0) 
print(X_train)
print(y_train)
0 Answers
Related