I am quite new to multiprocessing library and have question with regards to its Pool module when used with map(). Suppose I have 4 worker threads and 6 tasks to be completed. What I do is (using multiprocessing.dummy because I want to spawn threads and not processes)
from multiprocessing.dummy import Pool as ThreadPool
def print_it(num):
print num
def multi_threaded():
tasks = [1, 2, 3, 4, 5, 6]
pool = ThreadPool(4)
r = pool.map(print_it, tasks)
pool.close()
pool.join()
multi_threaded()
I want to understand how Pool.map() handles the tasks? Three options :
- Does it spawn 4 threads first, get the first 4 tasks complete and let the threads die. Then spawns 2 new threads for the remaining tasks?
- Does it spawn 4 threads, assign 4 tasks to them, as soon as some thread completes its task, assign new task to the same thread.
- Some other way.
This insight would be helpful as it will help me think of using Pool.map() more effectively in prod.