How to calculate the accuracy among two non-chronological order and different length lists using scikit-learn

Viewed 135

Background

Calculation of the matching rate of the two non-chronological order and different length lists

There is no meaning in indexes. Just want to know whether words in y_true are also in y_pred or not.

#answer
y_true = ['New York', 'city', 'hamburger', 'noisy']
#prediction
y_pred = ['beer', 'nostalgic', 'city', 'Berlin', 'sausage']

Problem

  1. How to acquire correct numbers for accuracy and other evaluation values

When comparing elements in the two lists, one pair (city-city) among four elements (y_true) and five elements (y_pred) is matching.

1/20 would be matched but the accuracy was 0.0.

  1. Dose padding make negative influence for the calculation?

Since it is essential to make arrays same length for calculating accuracy, precision, recall, and f1 score using sklear.

Are there other correct calculation for non-chronological order and different length lists?

Code

from sklearn.metrics import accuracy_score
from sklearn.metrics import precision_score
from sklearn.metrics import recall_score
from sklearn.metrics import f1_score
#answer
y_true = ['New York', 'city', 'hamburger', 'noisy']
#prediction
y_pred = ['beer', 'nostalgic', 'city', 'Berlin', 'sausage']

#padding
y_true = y_true + ['null']

print(y_true)
print(y_pred)

#evaliation
print(accuracy_score(y_true, y_pred))
print(precision_score(y_true, y_pred, average =  'micro'))
print(precision_score(y_true, y_pred, average =  'macro'))
print(precision_score(y_true, y_pred, average =  'weighted'))

print(recall_score(y_true, y_pred, average =  'micro'))
print(f1_score(y_true, y_pred, average =  'micro'))

#output
['New York', 'city', 'hamburger', 'noisy', 'null']
['beer', 'nostalgic', 'city', 'Berlin', 'sausage']
0.0
0.0
0.0
0.0
0.0
0.0
0 Answers
Related