Why are all of my cores used by concurrent.futures.ProcessPoolExecutor when fewer processes are requested to run?

Viewed 290

I am trying to understand why the following python script uses all of my CPU cores (8 core Ubuntu machine) when only 4 processes are requested

import os
import time
from concurrent import futures

def Sleeping(second):
    pid = os.getpid()
    print(f"Sleeping for {second} seconds (PID:{pid}).")
    time.sleep(second)

seconds = [15, 15, 15, 15]

with futures.ProcessPoolExecutor() as executor:
    for second in seconds:
        executor.submit(Sleeping, second)

Let's say the main script PID is 120. Then four PID (121, 122, 123, 124) appear on my system monitor for 15 seconds sleeping functions. However, extra PID (125, 126, 127, 128) also appear when it is obvious that they do nothing. It seems like a waste of resources. I do not have this issue when I am using multiprocessing.Process instead. Any insight on what is going on with concurrent.futures as opposed to multiprocessing is also appreciated.

1 Answers

We can gain some insight by looking at how many processes there are depending on the number of tasks. On a machine with 6 cores, and seconds = [15]*n I get the following number of processes...

  • n=0, processes=1
  • n=1, processes=4
  • n=2, processes=5
  • n=3, processes=6
  • n=4, processes=7
  • n=5, processes=8
  • n=6, processes=9
  • n=7, processes=9
  • n=8, processes=9
  • n=9, processes=9

Additionally, ThreadPoolExecutor takes an argument max_workers, which if I set I get the following number of processes...

  • n=9, max_workers=1, processes=4
  • n=9, max_workers=2, processes=5
  • n=9, max_workers=3, processes=6
  • n=9, max_workers=4, processes=7
  • n=9, max_workers=5, processes=8
  • n=9, max_workers=6, processes=9
  • n=9, max_workers=7, processes=9

So it looks like the process breakdown is: 1 main process, n-3 workers up to max_workers, and 2 other processes if at least one worker has started.

Moreover, if max_workers is not specified, it appears to be set to the number of cores. This explains why it maxes out at 9 = 1 + 6 + 2 processes in my case, since there are 6 workers, one per core.

So what are these 2 other processes? Well, this is an implementation detail, but it is most likely for process communication. It wouldn't be crazy to have a dedicated process for dispatching submissions and a dedicated process for receiving results. You can try to decipher it here: https://github.com/python/cpython/blob/3.10/Lib/concurrent/futures/process.py#L580 but I would just chalk it up to a fixed number of processes for communication and process management purposes. Depending on platform and some other factors like core count this fixed number could also vary, which could explain the difference between what I'm observing and what you're observing.

Try for yourself and see if you can reconcile the differences!

Related