I try to use cross-validation, and I don't know if this the right way or no I split the values into two-part
Then I use the X the value of the feature in PCA, then I used the output (features) from PCA in the cross validation function.
from sklearn.decomposition import PCA
from sklearn.ensemble import RandomForestClassifier
df = pd.read_csv('data.csv')
X = df.drop(['label'], axis = 1) Y = df['label']
pca = PCA(n_components=6) X_pca = pca.fit_transform(X)
model = RandomForestClassifier(n_estimators =400)
cv = StratifiedKFold(n_splits=5, random_state=123, shuffle=True)
n_scores = cross_val_score(model, pca, Y, scoring='accuracy', cv=cv,
n_jobs=-1, error_score='raise')
print('Accuracy: %.3f (%.3f)' % (mean(n_scores), std(n_scores)))
especially in this part :
n_scores = cross_val_score(model, pca, Y, scoring='accuracy', cv=cv,n_jobs=-1, error_score='raise')
the (pca) and (y) parameters is it in the right place?