In Python3, I have a starting dataframe in the format of a multilabel binary data:
df1:
"a" "b" "c" "d" "e"
1 1 0 0 1
0 0 1 0 1
1 0 0 0 0
0 1 1 0 1
What I need to achieve is this:
df2:
"a" "b" "c" "d" "e" "labels"
1 1 0 0 1 ["a", "b", "e"]
0 0 1 0 1 ["c", "e"]
1 0 0 0 0 ["a"]
0 1 1 0 1 ["b", "c", "e"]
To start, I tried using the inverse_transform() function from MultiLabelBinarizer from sklearn based on this previous stack question.
from sklearn.preprocessing import MultiLabelBinarizer
mlb = MultiLabelBinarizer()
mlb.fit(df1.columns)
mlb.inverse_transform(df1.values)
ValueError: Expected indicator for 15 classes, but got 5
I tried following the exact documentation from sklearn, but I am not sure where I went wrong. I tried tweaking a few of the parameters, but I do not understand what the issue is.
