Understanding the assumptions of different SHAP Methods

Viewed 212

I am interested in applying SHAP values to some work I am doing in machine learning, and notice that there are a number of different methods one can pick from on the github page: https://github.com/slundberg/shap

I have a neural network model, so as far as I understand I can use the functions: DeepExplainer, GradientExplainer or KernelExplainer. I am aware that DeepExplainer is based on DeepLift, and GradientExplainer is based on integrated gradients, but I am really struggling to find a clear outline of the assumptions each method makes. Is anyone able to clear up what assumptions each of these methods make, or point me in the direction of a source for this?

To be clear, I am not talking about the specific speed of each, I am considering the following points: Do any of them assume some property of the model? Do they assume inputs are independent? Which are appropriate for a mix of one-hot encoded and continuous variables? In short I cannot find a clear reference for this, or even a clear reference for what the GradientExplainer algorithm is even doing.

0 Answers
Related