Index mapping:
{
"settings": {
"index": {
"number_of_shards": 5,
"number_of_replicas": 2
}
},
"mappings": {
"properties": {
"vector": {
"type": "dense_vector",
"dims": 512
},
"category": {
"type": "keyword"
},
"name": {
"type": "keyword"
},
"source": {
"type": "keyword"
}
}
}
}
And I have around 600k documents in this index. I then use a query like that to retrieve similar documents based on vector:
{
script_score: {
query: { match_all: {} },
script: {
source: '(1.0 + cosineSimilarity(params.query_vector, \'vector\'))',
params: {
query_vector: vectorArray
}
},
min_score: 0
}
}
However, even with 5 shards Elastic takes around 1.3 seconds to complete the request. According to search profiler in Kibana, match phase takes 99.3% of that time.
Is there something I can do to improve my search performance and speed up document matching without changing script_score query?