I'm experimenting with Python's multiprocessing.shared_memory and multiprocessing.managers.SharedMemoryManager. I've created the following minimal code to get a grip of the process:
from multiprocessing import shared_memory
from multiprocessing import Pool
from multiprocessing.managers import SharedMemoryManager
import numpy as np
def main():
x = np.random.randn(10_000_000)
n_jobs = 25
n_workers = 4
with SharedMemoryManager() as smm, Pool(n_workers) as pool:
shm = smm.SharedMemory(x.nbytes)
y = np.ndarray(x.shape, dtype=x.dtype, buffer=shm.buf)
y[:] = x[:]
sample_means = pool.starmap(func, [(shm.name, x.shape, x.dtype)] * n_jobs)
def func(shm_name, shape, dtype):
shm = shared_memory.SharedMemory(name=shm_name, create=False)
x = np.ndarray(shape, dtype, buffer=shm.buf)
sample_mean = np.random.choice(x, size=10_000, replace=False).mean()
return sample_mean
if __name__ == '__main__':
main()
So using the shared_memory functionality we avoid making duplicates of the large numpy array x, which is good. However, it seems to me that the line y[:] = x[:] is actually creating a copy of x (in y). True, it happens only once (not for every process) but my question is - is it avoidable?
If x is a huge dataset I want to share between my workers I'd rather not make copies of it at all. Is it possible?