I am new to machine learning and currently working on a project with imbalance data. I want to balance the data using random undersampling. I am confused if i should do the undersampling after test train split or should i do undersampling 1st and then do train test split?
My approach : 1. I used train test split to get : X_train, y_train for training and X_test and y_test for testing. 2. I combined X_train and y_train into one data set and did the undersampling. 3. After undersampling, i performed Cross validation and model selection based on F1 score and using X_test.,Y_test for prediction.
Is my approach correct? Please correct me if i am wrong.