I have a setup where I run multiple (3) celery workers, and I have 8 different tasks: - celery - high frequency job: task 1, task 2 - low frequency jobs: task 3-8 each in their own kubernetes pod.
I want to implement monitoring using prometheus. For that, I'm using the library prometheus_client.
from celery import Celery, signals
from prometheus_client import start_http_server as start_prometheus_http_server
REDIS_HOST = os.environ.get("REDIS_HOST", "localhost")
BROKER_URL = f"redis://{REDIS_HOST}:6379/0"
app = Celery("tasks", broker=BROKER_URL)
app.conf.task_routes = {
"hifreq.main": {"queue": "main_queue"},
"hifreq.final": {"queue": "final_queue"},
"lowfreq.*": {"queue": "lowfreq_queue"},
}
@signals.celeryd_after_setup.connect
def setup_direct_queue(sender, instance, **kwargs):
start_prometheus_http_server(9090)
@app.task(name="hifreq.main")
def long_running_task():
data_loading()
DATA_LOADING_TIME = Summary(
"data_loading_seconds",
"Time spent loading the data",
)
@DATA_LOADING_TIME.time()
def data_loading():
pass
This starts the prometheus server (I think it starts one per worker). I've exposed it through an ingress/service such that I can reach the server, and when I navigate to the pod running the 'hifreq' worker, I get:
# HELP python_gc_objects_collected_total Objects collected during gc
# TYPE python_gc_objects_collected_total counter
python_gc_objects_collected_total{generation="0"} 5828.0
python_gc_objects_collected_total{generation="1"} 1643.0
python_gc_objects_collected_total{generation="2"} 294.0
# HELP python_gc_objects_uncollectable_total Uncollectable object found during GC
# TYPE python_gc_objects_uncollectable_total counter
python_gc_objects_uncollectable_total{generation="0"} 0.0
python_gc_objects_uncollectable_total{generation="1"} 0.0
python_gc_objects_uncollectable_total{generation="2"} 0.0
# HELP python_gc_collections_total Number of times this generation was collected
# TYPE python_gc_collections_total counter
python_gc_collections_total{generation="0"} 152.0
python_gc_collections_total{generation="1"} 13.0
python_gc_collections_total{generation="2"} 2.0
# HELP python_info Python platform information
# TYPE python_info gauge
python_info{implementation="CPython",major="3",minor="6",patchlevel="9",version="3.6.9"} 1.0
# HELP process_virtual_memory_bytes Virtual memory size in bytes.
# TYPE process_virtual_memory_bytes gauge
process_virtual_memory_bytes 3.19164416e+08
# HELP process_resident_memory_bytes Resident memory size in bytes.
# TYPE process_resident_memory_bytes gauge
process_resident_memory_bytes 4.4453888e+07
# HELP process_start_time_seconds Start time of the process since unix epoch in seconds.
# TYPE process_start_time_seconds gauge
process_start_time_seconds 1.58876788561e+09
# HELP process_cpu_seconds_total Total user and system CPU time spent in seconds.
# TYPE process_cpu_seconds_total counter
process_cpu_seconds_total 2.1
# HELP process_open_fds Number of open file descriptors.
# TYPE process_open_fds gauge
process_open_fds 30.0
# HELP process_max_fds Maximum number of open file descriptors.
# TYPE process_max_fds gauge
process_max_fds 1.048576e+06
which are the default Python metrics, but not the expected data_loading_seconds metric that I defined myself. I'm suspecting something is going wrong with the multiple workers each having their own server, but I'm not too sure what exactly is going wrong. Any help is appreciated!