How to use cross validation validation after train/test split

Viewed 47

I have used K-cross validation on the train set after splitting the data into train and test. But this gives an error which I think is due to indexing after the train and test split. Below is the code I used. How do I reset index after the train/train split or any other suggestions to deal with this error would be greatly appreciated. I have already tried df.reset_index() but this gives an error AttributeError: 'numpy.ndarray' object has no attribute 'reset_index'. Thank you.

X_train, X_test, y_train, y_test = train_test_split(X, y, test_size=0.1, random_state=99)

# k-fold cross validation
scores = list()
kfold = KFold(n_splits=10, shuffle=True)
# enumerate splits
for train_ix, test_ix in kfold.split(X_train):

    train_X, test_X = X_train[train_ix], X_train[test_ix]
    train_y, test_y = y_train[train_ix], y_train[test_ix]
    # fit model
    model = LinearRegression()
    model.fit(train_X, train_y)
    # evaluate model
    yhat = model.predict(test_X)
    score = np.sqrt(metrics.mean_absolute_error(yhat, test_y))
    print('Fold score : {}'.format(score))

KeyError: "Passing list-likes to .loc or [] with any missing labels is no longer supported. The following labels were missing: Int64Index([    3,     9,    10,    17,    19,\n            ...\n            41050, 41056, 41060, 41101, 41120],\n           dtype='int64', length=3708).
0 Answers
Related