Prometheus phantom counter resets

Viewed 204

I am trying to compute the rate of events published to a kafka topic. I have a metric kafka_topic_partition_current_offset that tracks the max offset for all my topics. This metric only ever increases, except if I were to delete and recreate the topic, which would start it over at 0.

The issue I'm having is that rate(kafka_topic_partition_current_offset[60s]) is producing massive spikes. I believe this is because it is detecting counter resets where there shouldn't be any.

Below, I've graphed resets(kafka_topic_partition_current_offset) next to kafka_topic_partition_current_offset with some scaling to bring them into perspective. These counter resets perfectly coincide with rate(..) spikes (not depicted).

enter image description here

My understanding is that counter resets are defined by when a counter decreases in value from one sample to the next, but we can clearly see that the yellow line never decreases, so I would expect no counter resets within the depicted interval.

My question is thus what other reasons would cause prometheus to detect resets?

More context:

  • my default scrape interval is 15s, but these kafka metrics are only scraped every 30s.
  • The resets don't disappear if I change the range vectors from [60s] to larger values
1 Answers

I see the same problem periodically. I would suggest that you look at

  1. your actual counter values
  2. the timestamps for your counter datapoints, and
  3. the result of resets() over your range

I suspect you'll see that your values and timestamps look fine, but resets() claims there's been reset. That's what I see when this happens.

I know, not an answer, but it's a starting place for looking into what appears to be a bug.

Related