Multiprocessing vs Threading in Python

Viewed 867

I am learning Multiprocessing and Threading in python to process and create large amount of files, the diagram is shown here diagram

Each of output file depends on the analysis of all input files.

Single processing of the program takes quite a long time, so I tried the following codes:

(a) multiprocessing

start = time.time()
process_count = cpu_count()
p = Pool(process_count)
for i in range(process_count):
    p.apply_async(my_read_process_and_write_func, args=(i,w))

p.close()
p.join()
end = time.time()

(b) threading

start = time.time()
thread_count = cpu_count()
thread_list = [] 

for i in range(0, thread_count):
    t = threading.Thread(target=my_read_process_and_write_func, args=(i,))
    thread_list.append(t)

for t in thread_list:
    t.start()

for t in thread_list:
    t.join()

end = time.time()

I am runing these codes using Python 3.6 on a Windows PC with 8 cores. However Multiprocessing method takes about the same time as the single-processing method, and Threading method takes about 75% of the single-processing method.

My questions are:

Are my codes correct?

Is there any better way/codes to improve the efficiency? Thanks!

3 Answers

Your processing is I/O bound, not CPU bound. As a result, the fact that you have multiple processes helps little. Each Python process in multiprocessing is stuck waiting for input or output while the CPU does nothing. Increasing the Pool size in multiprocessing should improve performance.

Follwing Tarik's answer, since my processing is I/O bound, I made serveral copies of input files, then each processing reads and processes different copy of these files. Now my codes run 8 times faster.

Now my processing diagram looks like this. multiprocessing My input files include one index file (about 400MB) and 100 other files(each size=330MB, can be considered as a file pool). In order to generate one output file, index file and all flles within the file pool need to be read. (e.g. First line of index file is 15, then line 15 of each files within the file pool need to be read to generate output file1.) Previously I tried multiprocessing and Threading without making copies, the codes were very slow. Then I optimized the codes by making copies of only the index file for each processing, so each processing reads copies of index file individually, and then reads the file pool to generate the output files. Currently, with 8 cpu cores, multiprocessing with poolsize=8 takes least time.

Related