I am trying to setup a new mapping for an index. Which is going to support partial keyword search and auto-complete requests powered by ES.
edgeNGram token filter with whitespace tokeniser seems a way to go. Till now my setting looks something like this:
curl -XPUT 'localhost:9200/test_ngram_2?pretty' -H 'Content-Type: application/json' -d'{
"settings": {
"index": {
"analysis": {
"analyzer": {
"customNgram": {
"type": "custom",
"tokenizer": "whitespace",
"filter": ["lowercase", "customNgram"]
}
},
"filter": {
"customNgram": {
"type": "edgeNGram",
"min_gram": "3",
"max_gram": "18",
"side": "front"
}
}
}
}
}
}'
The problem is with Japanese words! Does NGrams work on japanese letters? For e.g.: 【11月13日13時まで、フォロー&RTで応募!】
There is no whitespace in this - The document is not searchable with partial keywords, is that expected?