I'm quite new in ML so I'm here to ask some advices.
I'm doing a multilabel classification and my data is very unbalanced.
So I'm looking for a method to perform stratified sampling while splitting and found few methods for multilabel problem.
multi-label stratification with skmultilearn
from skmultilearn.model_selection import iterative_train_test_split X_train, y_train, X_test, y_test = iterative_train_test_split(X, y, test_size = 0.3)-
from iterstrat.ml_stratifiers import MultilabelStratifiedKFold mskf = MultilabelStratifiedKFold(n_splits=2, shuffle=True, random_state=0) for train_index, test_index in mskf.split(X, y): print("TRAIN:", train_index, "TEST:", test_index) X_train, X_test = X[train_index], X[test_index] y_train, y_test = y[train_index], y[test_index]
However I got errors in both cases :
- error for iterative_train_test_split
ValueError: not enough values to unpack (expected 2, got 1)
- error for iterative-stratification
ValueError: Supported target type is: multilabel-indicator. Got 'binary' instead.
I claned by data, each label has at least 4 samples. I don't really understand where these erros come from.
I suppose that these algorithms consider each label combinaison, not each label, as a individual so maybe some unique combinaisons give this error ? (But I still don't understand the second error)
Here is how my dataset looks like (title and description are features and tag is label) :
ID | title | description | tag
---------------------------------------------
000 | AAA | AAAAAAA | (C1, C2)
001 | BBB | BBBBBBB | (C8)
... ... ... ...
005 | TTT | TTTTTTT | (C1, C20, C500)