Process getting stuck after being launched from another process

Viewed 251

I was working on a specific type of application in Dash which required the action executed by pressing the button to be performed in a separate process. This process, in turn, was parallelizable, and in some cases spawned child processes for the efficient computation. The configuration given in this cases makes the child processes to get stuck. The code below reproduces the situation described as follows:

import multiprocessing
import time
import dash
from dash import html
from dash.dependencies import Input, Output

app = dash.Dash(__name__)

app.layout = html.Div([
    html.Button(id='refresh-button', children='Button'),
    html.Div(id='dynamic-container1')
])


def run_function(i):
    print('hello')
    time.sleep(15)
    print(f'hello world {i}')


def run_process():
    num = 1
    print('hello world 00000')
    process = multiprocessing.Process(target=run_function, args=(num,))
    process.start()
    process.join()
    print('hello world')


@app.callback(Output('dynamic-container1', 'children'), Input('refresh-button', 'n_clicks'))
def refresh_state(click):
    if click == 0 or click is None:
        return None
    p = multiprocessing.Process(target=run_process)
    p.start()
    p.join()

    return None

if __name__ == '__main__':
    app.run_server(debug=True)

The output of this application when pressing the button is always the following:

Connected to pydev debugger (build 172.3968.37)
Dash is running on http://127.0.0.1:8050/

 * Serving Flask app 'main' (lazy loading)
 * Environment: production
   WARNING: This is a development server. Do not use it in a production deployment.
   Use a production WSGI server instead.
 * Debug mode: on
pydev debugger: process 6884 is connecting

hello world 00000

which means that the first process utilizing the function run_process() was launched, however, the child process run_function(i) was not even started. I was trying to find an explanation in popular books on multiprocessing in Python and any guidance to these "chaining" processes, but to no avail. From my understanding, the new child process run_function(i) should occupy a separate core (if there is any free core) and not to depend on the resources consumed by the parent process run_process(). Could you, please, explain to me the mechanics of this? I have a doubt that in this code run_function(i) might be coerced to consume the same resources as run_process() does, so the system basically is just restricting any new process from starting from the same resources, but I would like to confirm it from more expert users of Python. I used Python 3.7 and Pycharm Community 2017.2.3 on Win7 to reproduce this example

1 Answers

I can not reproduce your error and i think this is because you are using windows ( i am on Linux with python 3.9), so i can not find the error for you, but maybe i can give you some hints:

First: To find the error, try to reduce the code to the core of the problem (you can remove the whole dash stuff to check if this is the error). In my tests the results were the same, with and without dash

Second: Windows and Linux handle the multiprocessing stuff a bit different: windows spawns the process:

The parent process starts a fresh python interpreter process. The child process will only inherit those resources necessary to run the process object’s run() method. In particular, unnecessary file descriptors and handles from the parent process will not be inherited. Starting a process using this method is rather slow compared to using fork or forkserver.

Unix Systems fork the process

The parent process uses os.fork() to fork the Python interpreter. The child process, when it begins, is effectively identical to the parent process. All resources of the parent are inherited by the child process. Note that safely forking a multithreaded process is problematic.

With Unix System you can spawn (multiprocessing.set_start_method("spawn")), with this i can't reproduce you error, but with this fork/spawn example i wanted to make clear, that there sometimes things are a bit different between windows and linux, even if the same packages are used. I think your understanding of multiprocessing is correct. (Maybe this site helps too.)

Third: In the docs are some programming guidelines you should be aware of. Maybe they will help, too. And in general, the multiprocessing package does not work well with a lot of interactive python shells (like IDLE or pycharm), this can may be better on newer versions. Maybe you should try it from terminal to check if this changes something.

I hope this helps a litle bit.

Related