Building a Pandas dataframe using multiprocessing leads to errors

Viewed 375

I am trying to build a big dataframe using a function that takes some arguments and return partial dataframes. I have a long list of documents to parse and extract relevant information that will go the big dataframe, and I am trying to do it using multiprocessing.Pool to make the process go faster.

My code looks like this :

    from multiprocessing import *
    from settings import *
    
    server_url = settings.SERVER_URL_UAT
    username =   settings.USERNAME
    password =   settings.PASSWORD
    
    def wrapped_visit_detail(args): 
        
        global server_url
        global username
        global password
        
        # visit_detail return a dataframe after consuming args
    
        return visit_detail(server_url, args, username, password) 
    
    # Trying to pass a list of arguments to wapped_visit_detail
    
    visits_url = [doc1, doc2, doc3, doc4]
    
    df = pd.DataFrame()
    pool = Pool(cpu_count())
    df = pd.concat( [ df,
                      pool.map( wrapped_visit_detail,
                                visits_url
                                )
                      ],
                    ignore_index = True
                    )

When I run this, I got this error

multiprocessing.pool.MaybeEncodingError: Error sending result: '<multiprocessing.pool.ExceptionWithTraceback object at 0x7f2c88a43208>'. Reason: 'TypeError("can't pickle _thread._local objects",)'

EDIT

To illustrate my problem I created this simple figure

enter image description here

This is painfully slow and not scalable at all

And I am looking to make the code not serial but rather as parallelized as possible

enter image description here

Thank you all for your great comments so far, yes, I a using shared variable as parameters to this function that pulls the files and extract the individuals dataframes, it seems ot be my issue indeed

I am suspecting something wrong in the way I call pool.map()

Any tip would be really welcome

1 Answers

You may have realised on your own, that there is actually no "sharing" possible, among Python-interpreter and its sub-processes. No sharing, only absolutely independent and "disconnected" replicas of the __main__'s original, (in Windows a complete, top-down) state-full copy of the Python-interpreter, with all its internal variables, modules and whatsoever. Any change of there copies is not propagated back into the __main__ or elsewhere. Once more, if you try to compose "Big dataframe" from ~ +100k individually pre-produced "partial dataframes", you will get an awfully if not unacceptably low performance.

Losing advantages from partial-producers' latency-masking plus headbanging into the RAM-allocation costs and potentially even a need to turn memory-I/O many times worse, falling into a trap of 10,000x slower physical/virtual-memory swapping, as in-RAM capacities ceased to be able to hold all data - as might happen upon ex-post attempt to
pd.concat( _a_HUGE_in_RAM_store_first_LIST_of_ALL_100k_plus_partial_DFs_, ... )
which is (unless you test Stack Overflow sponsors of Knowledge sense of humour and patience of others)
a no go ANTI-pattern.


Q :
" Any tip ... "

A :
Welcome to the realms of , no matter how simple this one is, here, due to Python-interpreted, process-to-process communication constraints

( unhandled EXC reporter says) :

multiprocessing.
           pool.MaybeEncodingError:

Error sending result:
'<multiprocessing.pool.ExceptionWithTraceback object at 0x7f2c88a43208>'.

Reason:
'TypeError("can't pickle _thread._local objects",)'

So,
here we are. Python-interpreter has to use "pickle"-like SER/DES whenever it tries to send/receive as single bit of data from one process ( typ. the __main__ upon sending launch parameters, or worker-processes upon sending their remote results' objects back to the __main__ ) to another.

That's fine, whenever the SER/DES-serialisation is possible ( here not being the case )

Options :

  • Best avoid any and all object-passing ( objects are prone to SER/DES-failures )
  • If obsessed with passing them, try some more capable SER/DES-encoder, replacing a default pickle with import dill as pickle has saved me many times ( yet not in every case, see above )
  • Design better problem-solving strategy, that does not rely on ill-assumed "shared"-use of Pandas dataframe amongst more (fully independent) sub-processes, that cannot and do not access "the same" dataframe, but theirs locally-isolated replica thereof (re-read more about sub-process independence and memory-space separation)

Nota bene:
School-book SLOC-s are nice in school-book sized examples, yet are awfully expensive performance ANTI-patterns in real-world use cases, the more in production-grade code-execution.

Tip:

If performance is The Target,
produce independent file-based results and join them afterwards. "Big dataframe" composition from "partial dataframes" represent a mix of sins, that cause you many performance ANTI-pattern problems, the failure to SER/DES-pickle being just a one, a small one (visible to naked eye)

Decisions ( & the costs associated with making them ) are in all cases on you.

You might like some further reads on how to boost performance.

Related