Can I set a variable amount of memory for SLURM jobs?

Viewed 230

I am on a school server, where we have a login node, and several GPU nodes that can be accessed with SLURM. Specifically, I am running several jobs on a machine with 512 GB RAM, 64 CPUS, and 8 GPUs.

I want to train a large deep-learning model with a ton of text. After some trial and error, it seems I need 100 GB of RAM to load all the text data (12 GB batch file on disk) successfully, before training starts. However, while training, it only uses about 30 GB of RAM. I want to run several versions of this model, so I can only do up to 5 before running out of RAM to allocate.

Is there a way to have SLURM use a variable amount of RAM? I don't want to hog up too many resources unnecessarily. This are the SBATCH directives I am currently using:

#SBATCH --mem-per-cpu=20G
#SBATCH --cpus-per-task=5
#SBATCH --nodelist=gpu-large
#SBATCH --gpus=1

I currently have 3 jobs running, and this is what I see from free:

              total        used        free      shared  buff/cache   available
Mem:          503Gi        94Gi       246Gi       1.6Gi       161Gi       403Gi
Swap:         4.0Gi        96Mi       3.9Gi

So my jobs indeed only are using 30 GB each. Help would be greatly appreciated!

1 Answers

Disclaimer: the following is from memory, I have no access to a slurm cluster ATM to check that.

You should be able to update the memory requirement of the job after it started. AFAIR, as a simple user (as opposed to a slurm admin or root), you can only lower the limits, not increase them (to prevent highjacking priority or ressources). Say, if your job has a memory-intensive step0 followed by a less intensive step1, you could try something like the following:

# Memory intensive part
step0
# Lower ressource reservation
scontrol update JobId=$SLURM_JOBID MinMemoryCPU=6G
# Less intensive step
step1

Some general info there: https://slurm.schedmd.com/scontrol.html

Related