I have been experimenting with the multiprocessing module in python, and I was wondering how the arguments of different parallelization methods are treated through the spawned processes. Here the code I used:
import os
import time
import multiprocessing
class StateClass:
def __init__(self):
self.state = 0
def __call__(self):
return f"I am {id(self)}: {self.state}"
CONTEXT = multiprocessing.get_context("fork")
nb_workers = 2
stato = StateClass()
def wrapped_work_function(a1, a2, sss, qqq):
time.sleep(a1 + 1)
if a1 == 0:
sss.state = 0
else:
sss.state = 123
for eee in a2:
time.sleep(a1 + 1)
sss.state += eee
print(
f"Worker {a1} in process {os.getpid()} (parent process {os.getppid()}): {eee}, {sss()}"
)
return sss
print("main", id(stato), stato)
manager = CONTEXT.Manager()
master_workers_queue = manager.Queue()
work_args_list = [
(
worker_index,
[iii for iii in range(4)],
stato,
master_workers_queue,
)
for worker_index in range(nb_workers)
]
pool = CONTEXT.Pool(nb_workers)
result = pool.starmap_async(wrapped_work_function, work_args_list)
pool.close()
pool.join()
print("Finish")
bullo = result.get(timeout=100)
bullo.append(stato)
for sss in bullo:
print(sss, id(sss), sss.state)
from which I get for example the following output:
main 140349939506416 <__main__.StateClass object at 0x7fa5c449dcf0>
Worker 0 in process 9075 (parent process 9047): 0, I am 140350069832528: 0
Worker 0 in process 9075 (parent process 9047): 1, I am 140350069832528: 1
Worker 1 in process 9077 (parent process 9047): 0, I am 140350069832528: 123
Worker 0 in process 9075 (parent process 9047): 2, I am 140350069832528: 3
Worker 0 in process 9075 (parent process 9047): 3, I am 140350069832528: 6
Worker 1 in process 9077 (parent process 9047): 1, I am 140350069832528: 124
Worker 1 in process 9077 (parent process 9047): 2, I am 140350069832528: 126
Worker 1 in process 9077 (parent process 9047): 3, I am 140350069832528: 129
Finish
<__main__.StateClass object at 0x7fa5c43ac190> 140349938516368 6
<__main__.StateClass object at 0x7fa5c43ac4c0> 140349938517184 129
<__main__.StateClass object at 0x7fa5c449dcf0> 140349939506416 0
The initial class instance stato has id 140349939506416, and keeps it through its lifetime as I would expect. Within the starmap_async method I get indeed two different instances of the same class (one for each worker/process), which I can modify and which retain their state property until the end of the script. Anyway the id of these instances is initially the same (140350069832528), and at the end of the script both of them have yet another id, which is also different from the one of the original instance.
Having the same id doesn´t mean that they have the same address in memory? How is it then possible that they retain a different state?
Is this behavior related to the fork context?