I've been exploring open-source libraries for Explaining Deep Learning model which is a multi-input Keras Sequential model trained on a dataset with text, numerical and categorical features. I am using TF v2.4.1 and the model network looks like this:
input1 = Input(shape=(200,)) # text input
emb = Embedding(vocab_size+1, 64)(input1)
lstm = Bidirectional(LSTM(64, dropout=0.2, recurrent_dropout = 0.1))(emb)
input2 = Input(shape=(100,)) # numerical and categorical
clf1 = Dense(64, activation='relu')(input2)
dropout = Dropout(0.2)(clf1)
clf2 = Dense(8, activation='relu')(dropout)
concat = concatenate([lstm, clf2])
output = Dense(3, activation='softmax')(concat)
model = Model(inputs=[input1 ,input2], outputs=[output])
Initially I tried getting explanations using DeepSHAP but found out that SHAP DeepExplainer has few incompatibility issues with TF v2, necessitating the use of tf.compat.v1.disable_v2_behavior(). Disabling the v2 behaviour affects the performance of my model, thus I'm searching for an alternative method that can provide explanations without affecting the model's performance.
And GradientSHAP is not working with model with embedding layer as described in SHAP [issue-1039[(https://github.com/slundberg/shap/issues/1039). So I want to apply Gradient SHAP for intermediate layers. Because my knowledge of TF is limited, I was unable to apply Graident SHAP to intermediate layers.
I can create the code below using the example given in the SHAP repo and the docstring:
gshap = shap.GradientExplainer(model = ([model.layers[3].output,model.layers[2].input],
concat_model.layers[-1].output ),
data = [text_input,tabular_input])
But the docstring says for TFv2 I should pass a TensorFlow function, not a tuple of input/output tensors.
Could someone help me in understanding what kind of tf function should be passed?