multilabel stratified split for unbalanced data

Viewed 262

I'm quite new in ML so I'm here to ask some advices.

I'm doing a multilabel classification and my data is very unbalanced.

So I'm looking for a method to perform stratified sampling while splitting and found few methods for multilabel problem.

  1. multi-label stratification with skmultilearn

    from skmultilearn.model_selection import iterative_train_test_split
    X_train, y_train, X_test, y_test = iterative_train_test_split(X, y, test_size = 0.3)
    
  2. iterative-stratification

     from iterstrat.ml_stratifiers import MultilabelStratifiedKFold
     mskf = MultilabelStratifiedKFold(n_splits=2, shuffle=True, random_state=0)
     for train_index, test_index in mskf.split(X, y):
         print("TRAIN:", train_index, "TEST:", test_index)
         X_train, X_test = X[train_index], X[test_index]
         y_train, y_test = y[train_index], y[test_index]
    

However I got errors in both cases :

  1. error for iterative_train_test_split

ValueError: not enough values to unpack (expected 2, got 1)

  1. error for iterative-stratification

ValueError: Supported target type is: multilabel-indicator. Got 'binary' instead.

I claned by data, each label has at least 4 samples. I don't really understand where these erros come from.

I suppose that these algorithms consider each label combinaison, not each label, as a individual so maybe some unique combinaisons give this error ? (But I still don't understand the second error)

Here is how my dataset looks like (title and description are features and tag is label) :

ID   |  title   |   description  |  tag
---------------------------------------------
000  |  AAA     |   AAAAAAA      |  (C1, C2)   
001  |  BBB     |   BBBBBBB      |  (C8) 
...     ...           ...            ...
005  |  TTT     |   TTTTTTT      |  (C1, C20, C500) 
0 Answers
Related