I'm looking to evaluate test performance of a random forest regressor in Python and, in addition to running cross-validation on the training-set, am wondering if it is appropriate to run some sort of correlation analysis between the predicted Y test results and the actual Y test results?
My possibly oversimplified thinking being that a significant correlation between the two would indicate that the predicted Y's are aligned with the actual test Y's and, as such, predictions are good...
Any alternative suggestions are more than welcome. Thanks.