How to find accuracy score on the test.csv after model being evaluated on the training data?

Viewed 43

To find the accuracy score, we execute model.score(X_train, y_train) for training set. and model.score(X_val, y_val) for validation set. Now, in my case, test data is a separate csv file. I have applied models on my training and test data. I know the score of training data but could not find the score on test data.

Below is my code:

model_dt = make_pipeline(
    SimpleImputer(strategy="mean"),
    DecisionTreeClassifier(random_state=42)
)
model_dt.fit(X_train, y_train)
acc_train = model_dt.score(X_train, y_train)
acc_val = model_dt.score(X_val, y_val)
print("reg model", acc_train, acc_val)
predictions_dt_reg = model_dt.predict(test)

**Note: After the above step I want to calculate the score on my test data **

1 Answers

So what can you do is call the test.csv and do the same data cleaning and transfromation steps on it. Next pass the cleaned x_test data to the model.predict().
It will give you predicted values/classes as per your problem. Next call this function this will help you get your accuracy only if you are dealing with classification problem:-

from sklearn.metrics import accuracy_score,confusion_matrix
print(accuracy_score(y_test,y_pred))  
print(confusion_matrix(y_test,y_pred)
#y_pred is the name of list in which xtest outpust are saved

If you are dealing with a regression problem then you can use MSE or RMSE to get the accuracy

from sklearn.metrics import mean_squared_error 
print(mean_squared_error(y_test,y_pred))
#y_pred is the output your model predicted with x_test data
Related