Ray: setting memory limit - workarounds

Viewed 30

Currently ray start ignores --memory settings, treating them like a burstable memory request, rather than a hard limit. Are there any known workarounds to cap memory usage of Ray Core servers?


More info

Currently parameters used to control all other resources [1] act as hard limits, but memory argument (of ray.init()) and --memory switch (of ray start) are definitely not hard limits. This becomes apparent when a container with Ray server is run separately from Jupyter Notebook server with the python client app, which allows us to distinguish between the client and server-side memory usage.

[1] i.e. object store memory and CPU / GPU numbers (provisioned with object_store_memory / --object-store-memory, num_cpus / --num-cpus and num_gpus / --num-gpus).

1 Answers

Unfortunately there is no way to do this at the moment, but it is an active area of development.

In the meantime, the best way to enforce a hard cap is to use a container with a memory limit, but note that this will kill the Ray node if it reaches the hard cap.

If the memory usage is mainly coming from the Ray application, the best way to work around this at the moment is to request more CPUs per task, to ensure that fewer of them run in parallel. If you're seeing unusually high memory usage from Ray system-level processes like the GCS or the raylet, then it may be a bug in Ray core.

Related