I have a list of lists that contain classificatory labels for a certain domain. Example:
data = [
['polmone', 'linfonodi'],
['osso'],
['polmone'],
['linfonodi', 'osso', 'polmone'],
['peritoneo', 'osso'],
['fegato'],
['polmone', 'linfonodi'],
['osso'],
['osso', 'fegato'],
]
The list has 331 lists and each of them can contain one or all the possible labels. The number of possible labels is 20.
I need to feed the list of lists of labels to a sklearn.neighbors.KNeighborsClassifier and was thinking of converting each possible label to a number (e.g. 0-19).
I was wondering about a most efficient way to perform this conversion.
I guess the 'stupid' way could be that of creating a dictionary with each unique label and the corresponding value, as in:
{'polmone': 0, 'linfonodi': 1, ..., 'label_19': 19}
...and then iterate over each element of the list and perform a str.replace().
I feel there should be a more efficient solution. Do you advice any?
Thanks in advance.
P.S. I searched for a similar topic, but couldn't find one. If I mistakenly didn't notice it, feel free to close this thread and send me to hell.
Edit:
First of all, I'd like to thank everyone for their answers, as every one of them has come to help for different issues I was encountering and I will encounter.
Now I want to share another solution I just found when dealing with KNeighborClassifier and a multiple-output target.
By feeding the encoded labels (both as strings or as integers, and both as simple lists or as numpy arrays), I had the following error:
Traceback (most recent call last):
File "embedding_gensim.py", line 111, in <module>
neigh.fit(doc_train, labls_train)
File "/home/matteo/anaconda3/envs/deep_l/lib/python3.7/site-packages/sklearn/neighbors/base.py", line 906, in fit
check_classification_targets(y)
File "/home/matteo/anaconda3/envs/deep_l/lib/python3.7/site-packages/sklearn/utils/multiclass.py", line 169, in check_classification_targets
raise ValueError("Unknown label type: %r" % y_type)
ValueError: Unknown label type: 'unknown'
I found that MultiLabelBinarizer solves the problem of feeding the classifier with a multi-label list of lists (or numpy arrays).
So, following @Alexander Rossa's solution:
binarized_labels = MultiLabelBinarizer().fit_transform(encoded_labels_list)
binarized_labels then is like:
[0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 1 0 0 0 0 0]
[0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0 0 0 0 0]
[0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 0 1 0 0 0 0 0]
[0 0 0 0 0 0 0 1 0 0 0 0 1 0 0 0 1 0 0 0 0 0]
...
The MultiLabelBinarizer() actually works directly with the lists of strings in split_labels .
Perhaps I am tackling the problem from the wrong perspective.