I built this inverted index:
{
'experiment': {'d1': [1, [0]], ..., 'd30': [2, [12, 40]], ..., 'd123': [3, [11, 45, 67]], ...},
'studi': {'d1': [1, [1]], 'd2': [2, [0, 36]], ..., 'd207': [3, [19, 44, 59]], ...}
}
For example, the term experiment appears in document 1 one time at index zero, in document 30 two times at indices 12 and 40, etc. I am wondering how I could count the number of occurrences of each term in the dictionary based on a dictionary of queries that looks like this:
{
'q1' : ['similar', 'law', ..., 'speed', 'aircraft'],
'q2' : ['structur', 'aeroelast', ..., 'speed', 'aircraft'],
...
'q225': ['design', 'factor', ..., 'number', '5']
}
The desired output would look something like this:
{
'q1' : ['d51', 'd874', ..., 'd717'],
'q2' : ['d51', 'd1147', ..., 'd14'],
...,
'q225': ['d1313', 'd996', ..., 'd193']
}
With keys representing the query and values representing the documents that the query appeared in, and the list would be sorted in descending order of total term frequencies