I have a function with list valued return values that I'm multiprocessing in Python and I need to concatenate them to 1D lists at the end. The following is a sample code for demonstration:
import numpy as np
import multiprocessing as mp
import random as rd
N = 4
L = list(range(0, N))
def F(x):
a = []
b = []
for t in range(0,2):
a.append('a'+str(t*x))
b.append('b'+str(t*x))
return a, b
pool = mp.Pool(mp.cpu_count())
a,b = zip(*pool.map(F, L))
pool.close()
print(a)
print(b)
A = np.concatenate(a)
B = np.concatenate(b)
print(A)
print(B)
The output for illustration is:
(['a0', 'a0'], ['a0', 'a1'], ['a0', 'a2'], ['a0', 'a3'])
(['b0', 'b0'], ['b0', 'b1'], ['b0', 'b2'], ['b0', 'b3'])
['a0' 'a0' 'a0' 'a1' 'a0' 'a2' 'a0' 'a3']
['b0' 'b0' 'b0' 'b1' 'b0' 'b2' 'b0' 'b3']
The problem is that the list L that I'm processing is pretty huge and that the concatenations at the end take a huge amount of time which minimizes the advantage over serial processing considerably.
Is there some clever way to avoid the concatenation or alternatively a faster method to perform the concatenation? I've been fiddling with queues but this seems kind of very slow.
Note: This seems to be a similar question as Add result from multiprocessing into array.