Problems understanding nDCG format in pytrec_eval?

Viewed 136

I am using pytrec_eval to calculate nDCG scores. For example, for qrel:

qrel = {
    'q1': {
        'd1': 0,
        'd2': 1,
        'd3': 0,
    }
}

And run:

run = {
    'q1': {
        'd1': 1.0,
        'd2': 0.0,
        'd3': 1.5,
    }
}

The nDCG score can be calculated like this:

import pytrec_eval
import json    
evaluator = pytrec_eval.RelevanceEvaluator(
    qrel, {'ndcg'})

print(json.dumps(evaluator.evaluate(run), indent=1))


 "q1": {
  "ndcg": 0.5
 }

My understanding is that nDCG takes into account the order of the retrieved documents indices, however, if you change the document ordering in the run you still get the same nDCG score, for example:

run2 = {
    'q1': {
        'd1': 1.0,
        'd3': 1.5,
        'd2': 0.0,

    }
}

evaluator = pytrec_eval.RelevanceEvaluator(qrel, {'ndcg'})
print(json.dumps(evaluator.evaluate(run2), indent=1))

Is this the expected behavior of calculating nDCG? What is the usage of the qrel? My undertanding is that the qrel tells you how relevant is the retrieved document, while run, is the resulting ranking of your query and IR system. Then, why if I change the order of run the nDCG score is the same?

0 Answers
Related