Prefix search with frequency count

Viewed 359

At the moment when I index a text, I store the frequency count of each word in a database. This works just fine since all searches are based on whole words and all possible searches are known. But now I want to add the option of a prefix search (search of a part of a word). I can get the results/hits from a prefix search with elasticsearch by using this:

GET /my_index/address/_search
{
    "query": {
        "prefix": {
            "main_text": "word_part"
        }
    }
}

see: https://www.elastic.co/guide/en/elasticsearch/guide/current/prefix-query.html

This is my current mapping:

{
    "my-index":{
        "mappings":{
            "doc":{
                "properties":{
                    "keycounter":{
                        "properties":{
                            "counter": {"type":"integer"},
                            "keyword":{"type":"keyword"}
                         }
                    },
                    "main_text":{
                        "type":"text", 
                        "fielddata":true
                    },
                    "main_text_keycounter":{
                        "properties":{
                            "counter":{
                                "type":"long"
                            },
                            "keyword":{
                                "type":"text", 
                                "fields":{
                                    "keyword":{
                                        "type":"keyword",
                                        "ignore_above":256
                                    }
                                }
                            }
                        }
                    },
                    "time_written":{
                        "type":"date"
                    },
                    "translated_text":{
                        "type":"text",
                        "fielddata":true
                    },
                }
            }
        }
    }
}

But I don´t want to count the frequency for each result I get since it will cost O(N) for each text. Is there some smart way of storing/getting frequency count from this type of a search using elasticsearch?

2 Answers

You can use the doc-termvectors feature of elasticsearch to get the term statistics and count of the terms. Like that you can store your document using the mapping and get the statistics of the prefix term when you query it. Of course this approach provides you with term statistics per result document so you'd have to aggregate it for all of your results.

Here is an example for a mapping, indexed document and a doc-termvectors query. You can also use a edge-ngram tokenizer to get statistics for prefix-terms.

Mapping:

PUT /my-index
{
  "mappings": {
    "doc": {
      "properties": {
        "main_text": {
          "type": "text",
          "fielddata": true,
          "term_vector": "with_positions_offsets_payloads",
          "store": true
        }
      }
    }
  }
}

Index document:

POST /my-index/doc/1
{
  "main_text": "foo bar foo"
}

Get termvectors:

POST /my-index/doc/1/_termvectors

Results:

...
"terms": {
    ...
    "foo": {
      "term_freq": 2,
      "tokens": [
        {
          "position": 0,
          "start_offset": 0,
          "end_offset": 3
        },
        {
          "position": 2,
          "start_offset": 8,
          "end_offset": 11
        }
      ]
    }
    ...

Edit

If you want to get the termvectors for mutiple documents you can use the _mtermvectors endpoint. It will provide you with the statistics for multiple documents. It will, however, not count term frequencies over all documents which is as I understand your question what you want. As a solution, you might store the results of termvectors in your elastic (either same index, or a separate) and then use an aggregation to count the overall term counts.

POST /my-index/doc/_mtermvectors
{
  "ids": [
    "1",
    "2"
  ],
  "parameters": {
    "fields": [
      "main_text"
    ],
    "term_statistics": true
  }
}

Edit

Then I think that the solution is to call termvectors for all documents and store the results, i.e. all the term and subterm frequencies in another index. By aggregating the results based on your search queries, wou'd get the results that you wish.

Take a look at this answer, suggesting use of finite state transducer to speed up prefix searches for completion suggester. Looks pretty neat and claimed to be equivalent to trie usage

Related