what is the right way to use cross validation with feature extraction

Viewed 68

I try to use cross-validation, and I don't know if this the right way or no I split the values into two-part

Then I use the X the value of the feature in PCA, then I used the output (features) from PCA in the cross validation function.

from sklearn.decomposition import PCA 
from sklearn.ensemble import RandomForestClassifier


df = pd.read_csv('data.csv')

X = df.drop(['label'], axis = 1) Y = df['label']

pca = PCA(n_components=6) X_pca = pca.fit_transform(X)


model = RandomForestClassifier(n_estimators =400)


cv = StratifiedKFold(n_splits=5, random_state=123, shuffle=True)
n_scores = cross_val_score(model, pca, Y, scoring='accuracy', cv=cv,
n_jobs=-1, error_score='raise')

print('Accuracy: %.3f (%.3f)' % (mean(n_scores), std(n_scores)))

especially in this part :

n_scores = cross_val_score(model, pca, Y, scoring='accuracy', cv=cv,n_jobs=-1, error_score='raise')

the (pca) and (y) parameters is it in the right place?

0 Answers
Related