I'm trying to query a field that contains a lot of words in it, while each word is already multiplied by it's dominance in the document. Therefore, term frequency is what I need here, while idf really changes the documents scoring.
For example, I have two documents that contain a 'words' field, the first doc has the word 'fashion' 1000 times, and the second doc has the word 'baby' 1000 times as well. Score evaluation by term frequency alone, would return the same score (approx.) for both, but the idf additional evaluation changes the results, which I would like to avoid.
The mapping I have is pretty straight forward:
"query": {
"match": {
"words": {
"query": "fashion baby"
}
}
}
And the query is a simple match query:
The explain plan looks like this:
{
"_shard": "[keywords_words][4]",
"_node": "QQC692ZKQIif-_2eo_wO6Q",
"_index": "keywords_words",
"_type": "topic_words",
"_id": "133",
"_score": 407.61816,
"_source": {},
"_explanation": {
"value": 407.61816,
"description": "sum of:",
"details": [
{
"value": 407.61816,
"description": "weight(words:baby in 0) [PerFieldSimilarity], result of:",
"details": [
{
"value": 407.61816,
"description": "score(doc=0,freq=1000.0), product of:",
"details": [
{
"value": 3.5902672,
"description": "queryWeight, product of:",
"details": [
{
"value": 3.5902672,
"description": "idf, computed as log((docCount+1)/(docFreq+1)) + 1 from:",
"details": [
{
"value": 2,
"description": "docFreq",
"details": []
},
{
"value": 39,
"description": "docCount",
"details": []
}
]
},
{
"value": 1,
"description": "queryNorm",
"details": []
}
]
},
{
"value": 113.53422,
"description": "fieldWeight in 0, product of:",
"details": [
{
"value": 31.622776,
"description": "tf(freq=1000.0), with freq of:",
"details": [
{
"value": 1000,
"description": "termFreq=1000.0",
"details": []
}
]
},
{
"value": 3.5902672,
"description": "idf, computed as log((docCount+1)/(docFreq+1)) + 1 from:",
"details": [
{
"value": 2,
"description": "docFreq",
"details": []
},
{
"value": 39,
"description": "docCount",
"details": []
}
]
},
{
"value": 1,
"description": "fieldNorm(doc=0)",
"details": []
}
]
}
]
}
]
}
]
}
},
{
"_shard": "[keywords_words][1]",
"_node": "QQC692ZKQIif-_2eo_wO6Q",
"_index": "keywords_words",
"_type": "topic_words",
"_id": "490",
"_score": 344.91177,
"_source": {},
"_explanation": {
"value": 344.9118,
"description": "sum of:",
"details": [
{
"value": 344.9118,
"description": "weight(words:fashion in 2) [PerFieldSimilarity], result of:",
"details": [
{
"value": 344.9118,
"description": "score(doc=2,freq=1000.0), product of:",
"details": [
{
"value": 3.3025851,
"description": "queryWeight, product of:",
"details": [
{
"value": 3.3025851,
"description": "idf, computed as log((docCount+1)/(docFreq+1)) + 1 from:",
"details": [
{
"value": 2,
"description": "docFreq",
"details": []
},
{
"value": 29,
"description": "docCount",
"details": []
}
]
},
{
"value": 1,
"description": "queryNorm",
"details": []
}
]
},
{
"value": 104.43691,
"description": "fieldWeight in 2, product of:",
"details": [
{
"value": 31.622776,
"description": "tf(freq=1000.0), with freq of:",
"details": [
{
"value": 1000,
"description": "termFreq=1000.0",
"details": []
}
]
},
{
"value": 3.3025851,
"description": "idf, computed as log((docCount+1)/(docFreq+1)) + 1 from:",
"details": [
{
"value": 2,
"description": "docFreq",
"details": []
},
{
"value": 29,
"description": "docCount",
"details": []
}
]
},
{
"value": 1,
"description": "fieldNorm(doc=2)",
"details": []
}
]
}
]
}
]
}
]
}
}
The results scores are pretty important, so constant_score query is unfortunately not relevant in this case. Is there any way to make the query use only the term frequency weight?
Thanks a lot in advance!