The Elasticsearch documentation mentions that a scroll query or point in time operation can put greater pressure on shard's disk, memory, or os via open file handles as older segments cannot be merged.
Is the amount of data retained due to an open search context that would otherwise be deleted proportional to the size of the segment the updated data happened to be on and not proportional to the amount of data that's updated as perceived by the client? For example, if a client updated a 5KB document and internally the data for this document was on a 10MB segment which ends up getting merged, the entire 10MB segment would be retained when it otherwise would've been deleted. So in essence, the memory/disk impact of this context staying open is 10MB rather than 5KB. Is this correct?
If this is the case, is there any bound on how large a retained segment can be? Would a faster rate of indexing or larger document being indexed result in seeing more memory consumption? Is there anyway to do some back of the envelope calculation based on an applications access patterns to determine what kind of worst case you might expect? - or will there always be some probability that an unlucky update causes a merging operation and a large segment to be retained that can potentially cause resource exhaustion?