Wait for a group of processes, without a leader?

Viewed 93

I have some code that forks/waits, but it also might end up using some third party code that may also fork/wait. To limit the amount of processes I fork, I want to wait for a process to exit, if too many have been forked already. If I wait for any process though, I might wait on a process that third party code then expects to be able to wait on, leaving that third party code with a failure result and no information on exit status. My own code will also not work right, since I'll end up with a negative amount of active processes, if I end up waiting for more processes than I fork.

I was going to try to keep my forking limited to a process group, so I could wait on that, but where do I get a special "my code" process group, to use in my blocking version of fork? I can't get third party code to set a special process group themselves, and I can't use any process group except for the pid of the process doing all these forks, which third party code will also use. I could use one of the child processes as the process group leader, but then when that child exits I'm hosed, since I'll have to wait on two process groups now, then three, and so on. Should I just realloc a growing array of process groups that still have child processes in them? I could fork a process that immediately exits, then use that "zombie" process as the process group leader, but then when I wait on any process in that group, it'll clean up the zombie process leaving me once again with no process group leader. I'd use setrusage to limit subprocesses, but then when fork fails from too many subprocesses, I have no way to wait for any of those subprocesses to exit before trying to fork again.

My best idea so far is a heap allocated growing list of lists of subprocesses, each with a possibly dead process group leader. Can you still wait on a process group if the leader has exited though? If the pids overflow and cycle around, and a new process happens to get that pid, will it just magically become the process group leader? Should I be using something with semaphores? Two processes with every fork, one to wait on the other then increment the semaphore? A heap allocated growing list of pids to wait for individually, just randomly guessing which pid will exit first? I have to keep my own custom "zombie process" table, right? So that I can "wait" for a process that's already been waited for and still get the exit status? Am I just forbidden from using third party code in any process that forks, and need to always use the code in child processes so the parent can't inadvertently wait on any internal forks?

1 Answers

What I ended up doing ...seems like it was effective. "No good solutions" etc. but what I did was:

  • process A forks process B, then process A just waits on B
  • process B sets its own process group to itself (B)
  • process B can have a special fork function then, that sets the process group to A after forking (A being the "grandparent" process)
  • any naive fork will just use the process group B
  • if this system uses itself, then B will fork C, and C's subprocesses will use B as a process group. So not even that will interfere with process group A
  • if B counts too many processes, it just waits on group A, to get any of (and only) the child processes that have been counted

One problem is that shells rely on process groups for killing a process tree. They won't kill any subprocesses that set a different process group. So I had to use the non-platform-specific prctl(PR_SET_PDEATHSIG, ...) to have subprocesses kill themselves when the parent process dies. And furthermore, because PDEATHSIG gives you the thread ID, not the process ID, I had to use PR_SET_CHILD_SUBREAPER on process B, so that anything getting a PDEATHSIG would get one for when B dies, but can ignore when the thread within B exits.

A platform independent way to do this might be just poll kill(getppid(), 0) before every fork, to see whether you should die rather than fork. Checking the return value of setpgid might work too, but I don't know if it forbids you from using process A as a process group, if A has died, the PID number has cycled around, and a totally unrelated process happened to get A's old PID.

Related