I'm relatively new to ElasticSearch, and working on getting some partial matches across multiple fields to work. Assume, for example, I have the following three documents indexed:
{
"document-id": "Patient1",
"document-type": "patients",
"firstName": "Benjamin",
"lastName": "Carlton",
"medicalRecordNumber": "111-222-3333"
}
{
"document-id": "Patient2",
"document-type": "patients",
"firstName": "Carly",
"lastName": "Benson",
"medicalRecordNumber": "111-222-3334"
}
{
"document-id": "Patient3",
"document-type": "patients",
"firstName": "Jason",
"lastName": "Benson",
"medicalRecordNumber": "111-222-3335"
}
I would like to design an analyzer and search query such that searching for:
- "ben" matches all three (that's easy)
- "ben carl" matches #1 and #2
- "carl ben" also matches #1 and #2
- "benj carl" matches only #1 (does not follow as naturally from the previous ones as I thought, given the way ngram tokenizers seem to work)
- "carlt ben" matches only #2 (same)
- "benj carlt" will have no matches
- "111-222-3334" matches only #2
I feel like I'm close, using the following analyzer:
{
"settings": {
"analysis": {
"tokenizer": {
"partialMatchTokenizer": {
"type": "edge_ngram",
"min_gram": 2,
"max_gram": 10
}
},
"analyzer": {
"partialMatchAnalyzer": {
"type": "custom",
"tokenizer": "partialMatchTokenizer",
"char_filter": [],
"filter": [
"lowercase"
]
}
}
}
},
"mappings": {
"_doc": {
"properties": {
"lastName": {
"type": "text",
"analyzer": "partialMatchAnalyzer"
},
"firstName": {
"type": "text",
"analyzer": "partialMatchAnalyzer"
}
}
}
}
}
And the following query:
{
"query": {
"multi_match": {
"query": "carlt ben",
"type": "cross_fields",
"fields": [
"firstName",
"lastName",
"medicalRecordNumber"
],
"operator": "or"
}
}
}
But it's not quite there. The "or" seems too permissive; an "and" seems too restrictive. And sometimes the n-gram matching seems to provide unexpected results. For example, the above query ("carlt ben") matches both #1 and #2 (i.e., "carlt" matches "Carly", presumably because the "carl" n-gram matches). Also, oddly, "carlt ben" and "ben carlt" provide two different result sets (#1 & #2 vs. #1 & #2 & #3).
Any ideas on how I need to change my analyzer and/or my query to get the results I described above?