I am getting the error in this line of code (X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)
Error: ValueError: Found input variables with inconsistent numbers of samples: [40000, 10000]. Seems like after the vectorization, the array size gets change and that does not match with y. Seeking support in resolving the error. Thanks in advance
Output: (10000, 4) (10000,) (40000, 1500) (10000,)
import pandas as pd
from sklearn.model_selection import train_test_split
from sklearn.feature_extraction.text import CountVectorizer
# Import dataset:
dataset = pd.read_excel(r"C:\Users\HPS1RT\Downloads\test\Safety_Prediction.xlsx", nrows=10000)
dataset[["Safety"]] *= 1
# Assign values to the X and y variables:
X = dataset.iloc[:, :-1].values
y = dataset.iloc[:, 4].values
print(X.shape)
print(y.shape)
#vectorization
vectorizer = CountVectorizer(max_features=1500, min_df=5, max_df=0.7)
X = vectorizer.fit_transform(X.ravel()).toarray()
from sklearn.feature_extraction.text import TfidfTransformer
tfidfconverter = TfidfTransformer()
X = tfidfconverter.fit_transform(X).toarray()
print(X.shape)
print(y.shape)
# Split dataset into random train and test subsets:
X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.2, random_state=0)
print(X_train)
print(y_train)