How to encode column of list for catboost?

Viewed 118

I have a dataset where some columns contain lists:

import pandas as pd

df = pd.DataFrame(
    {'var1': [1, 2, 3, 1, 2, 3],
     'var2': [1, 1, 1, 2, 2, 2],
     'var3': [["A", "B", "C"], ["A", "C"], None, ["A", "B"], ["C", "A"], ["D", "A"]]
    }
)

    var1    var2    var3
0      1       1    [A, B, C]
1      2       1    [A, C]
2      3       1    None
3      1       2    [A, B]
4      2       2    [C, A]
5      3       2    [D, A]

As the values within the lists of var3 can be shuffled and we can't assume any specific order the only way I can think of to prepare the columns for modelling is one-hot encoding. It could be done quite easily:

df["var3"] = df["var3"].apply(lambda x: [str(x)] if type(x) is not list else x)

mlb = MultiLabelBinarizer()
mlb.fit_transform(df["var3"])

resulting in:

array([[1, 1, 1, 0, 0],
   [1, 0, 1, 0, 0],
   [0, 0, 0, 0, 1],
   [1, 1, 0, 0, 0],
   [1, 0, 1, 0, 0],
   [1, 0, 0, 1, 0]])

However, quoting catboost documentation:

Attention. Do not use one-hot encoding during preprocessing. This affects both the training speed and the resulting quality.

Therefore, I'd like to ask if there's any other way I could encode this column for modelling with catboost?

0 Answers
Related