We are using Prometheus for scraping metrics and keep them for 8 hours (for longer peried we are using Victoria Metrics). I tried to find out, how much MEM should Prometheus have for current amount of series. I find out How much RAM does Prometheus 2.x need for cardinality and ingestion? with calculator. From the article, there is some points which are unclear to me:
- Number of Time Series* - it should be result of
max_over_time(prometheus_tsdb_head_series[1d])query. In my case, there is multiple results, so I sum it withsum(max_over_time(prometheus_tsdb_head_series[1d])). Is this approach right? - Average Labels Per Time Series* - count number of key-value pairs of all labels and count average. E.g.
{a="123", b="456"}is counted as 2. - Number of Unique Label Pairs* - find uniq key-value pairs. E.g.
{a="123", b="456"}and{a="123", b="abc"}equals to 2 uniq label pairs - Average Bytes per Label Pair* (Including the =, "" and ,) - e.g.
{ "__name__": "up", "app_version": "0509b54", "instance": "1", "job": "event-metric-exporterd", "stack_version": "MR" }. I remove spaces and{,}. So final string is"__name__":"up","app_version":"0509b54","instance":"1","job":"event-metric-exporterd","stack_version":"MR". Then, I encode it withutf-8and get length of string, in this case 106. I do this process for each uniq labels and count average. Is this right? - Bytes per Sample - there is query
rate(prometheus_tsdb_compaction_chunk_size_bytes_sum[1d])/rate(prometheus_tsdb_compaction_chunk_samples_sum[1d])which again, in my case return multiple series. It takemax()because it is worst case scenario.
The calculator counts:
- Cardinality Memory - how much memory Prometheus needs for various series/cardinality? Do I understand correctly?
- Ingestion Memory - memory used by Prometheus to ingest data - load and process it from targets?
Do I understand correctly Cardinality Memory and Ingestion Memory? Is my calculation process correct? Is there better way how to estimate memory for Prometheus?
Using Prometheus in docker prometheus:v2.29.1.