I have a RandomForest classification model trained with (30000 x 164) data (train-70%, test-30%).
**RandomForestClassifier(n_estimators= 200, max_features= 'sqrt', random_state=40)**
test results= Sensitivity - 75, Specificty - 99
Because I have imbalanced classes(1s-10%, 0s-90%), I had to get the probabilities and extract the classes based on a threshold value probability(0.25 in my case) instead of the default 0.5.
test results with 0.25 probability threshold = Sensitivity - 91, Specificty - 96
Now, I have a new inference dataset (168x164) I want to predict the classes for, using the saved RF model(mlflow). When I prepare this data and make sure the column order of the inference data match with X_train, the model (threshold=0.25) is predicting 97%-99.5% as a class '1' when I am expecting only 5-10% of the observations as class '1' and the rest as '0'.(the default predictions or for the threshold 0.5 - the model is predicting more than 50% of data as class '1', which is also terrible) I checked this for multiple inference datasets, I see the same problem.
When I do not match the column order of train and inference data, the results are more reasonable compared to the above case. I know the column order should match, but am not sure what is causing the issue. Can someone help me with this?
Please let me if you need more in ormation to understand the problem. Thank you in advance.