Shouldn't the variables ranking be the same for MLP and RF?

Viewed 58

I have a question about variable importance ranking. I built an MLP and an RF model using the same dataset with 34 variables and achieved the same accuracy on a similar test dataset. As you can see in the picture below the top variables for the SHAP summary plot and the RF VIM are quite different. Interestingly, I removed the low-ranked variable from the MLP and the accuracy increased. However, the RF result didn’t change. Does that mean the RF is not a good choice for modeling this dataset? It’s still strange to me that the rankings are so different: SHAP summary plot vs. RF VIM, I numbered the top and low-ranked variable

enter image description here

1 Answers

Shouldn't the variables ranking be the same for MLP and RF?

No. There may be tendency for different algos to rank certain features higher, but there is no reason for ranking to be the same.

Different algorithms:

  1. May have different objective functions to achieve intended goal.
  2. May use features differently to achieve min (max) of the objective function.

On top, what you cite as RF "feature importances" (mean decrease in Gini) is only one of the many ways to calculate "feature importance" for RF (including which metric you use, and how you calculate total decrease due to a feature). In contrast, SHAP is model agnostic when it comes to explaining feature contributions to outcome.

In sum:

  1. Different models will have different opinions about what is important and not. What is important for one algo may be not so important for another and vice versa. It doesn't tell anything about applicability of a model to a specific dataset.
  2. Use SHAP values (or any other feature importance metric that you and your clients understand) to explain a model (if necessary).
  3. Choose "best" model based on your goals: performance or explainability.
Related