I have an ES index which stores the unique key and last updated date for each document. I need to write an APi which will be used to sync the data related to this key, either delta (based on the date stored, e.g. give me data updated after 3rd Mar 2020)
Rough ES mapping:
{
"mappings": {
"userdata": {
"_all": {
"enabled": false
},
"properties": {
"userId": {
"type": "long"
},
"userUUID": {
"type": "keyword"
},
"uniqueKey":{
"type":"keyword"
},
"updatedTimestamp":{
"type":"date"
}
}
}
}
I will use this ES index to find the list of such unique keys matching the date filter and build the remaining details for each key from cassandra.
The API is stateless.
The no. of documents matching the date filter could be in thousands to few hundred thousand. Now, when synching such data, the client will need to paginate the results.
To paginate, I plan to use 'lastSynchedUniqueKey'. For each subsequent call, the client will provide this value and the API will internally perform a range query on this field and fetch the data with uniqueKey > lastSynchedUniqueKey
So, ES query will have following components:
- search query : (date rage query) + (uniqueKey > lastSynchedUniqueKey) + (query on username)
- sort : on uniqueKey in asc order
- size : 100 --> this is the max pageSize (suggest if it can be changed based on total no. of documents to be synced. Only concern being, don't want to load the ES cluster with these queries. There will be other indices in the cluster which are used for user-facing searches.)
What is better option to perform pagination in this case:
pagination: using (from + size) and filter and sort param: I know this will not performant.
scroll: with same filter and sort param
ES document suggests using '_doc' for sorting for scrolls. Which is not possible in my case. Is it ok to use a field in the index instead?
Is scroll faster than search_after?
Please provide your inputs about sorting and pagination from client perspective and internally.