Generate percentage prediction label by using BERT multi-label classification

Viewed 391

I'm currently working on multi-label classification task for text data. I have a dataframe with an ID column, text column and several columns which are text label containing only 1 or 0.

I used an existing solution proposed on this website Kaggle Toxic Comment Classification using Bert which permits to express in percentage its degree of belonging to each label.

Now, that I've trained my model I would like to use my model with new unlabeled text in order to obtain percentage of belonging to each label :

I've found this solution on this website and more especially this part code that I want to add at the end of my Kaggle code :

texts = [
  '.........',
  '.........',
  '..........',
  '..........',
]

for text in texts:
  ids, segments = tokenizer.encode(text, max_len=SEQ_LEN)
  inpu = np.array(ids).reshape([1, SEQ_LEN])
  predicted = (model.predict([inpu,np.zeros_like(inpu)]) >= 0.5).astype(int)
  labels = [
    label
    for i, label in enumerate(labels_ordered)
    if predicted[0][i]
  ]
  print ("%s: %s" % (text, labels))

But this solution only permit me to obtain class prediction but not the percentage percentage prediction for each class.

Do you have any idea how can I do to adapt this last pat of code to my Kaggle code and get percentage prediction ?

0 Answers
Related