I have a parallelised program using concurrent.futures/ThreadPoolExecutor:
from concurrent.futures import ThreadPoolExecutor as PoolExecutor
import numpy as np, timeit
start = timeit.default_timer()
n = 2
def f(samp):
t = samp ** 10
samps = np.random.uniform(low=0, high=1, size=(100000,))
with PoolExecutor(max_workers=n) as executor:
for _ in executor.map(f, samps):
pass
print(f"time: {timeit.default_timer() - start}")
It takes about 3s to run.
If I run it sequentially without parallelising, i.e.:
for samp in samps: t = samp ** 10
It takes about 0.05s to run (i.e. 100,000 iterations).
Why is the parallelised version taking so much longer. NB increasing max_workers also increases run time. Also, this maybe a silly code example but my original code was processing 800 files - it also took longer than the sequential version.