Running multiple ML models in Parallel in python

Viewed 476

I have a Flask endpoint that that takes an input string and runs 10 different ML pipelines on the same string, then combines and returns the results.

Each model takes 0.2 seconds to run. I am looking to parallelize this task.

So far I've tried the following:

  • Multiprocessing - The ML pipeline is not serializable. So I cant use this.
  • Joblib - Since the ML pipeline is fairly complicated, the subprocess creation overhead is too big and it takes longer to run compared to running the models sequentially. I have verified that I am able to run the models in parallel with this method, but the creation of the subprocesses takes a lot of time.
  • pyspark - Same problem as joblib. Parallel task takes longer to run.

The main issue is that Joblib creates new workers each time a new message arrives and loading the model and setting up the pipeline takes up a lot of time. I am looking for a framework that can setup workers, load the models in memory and just waits and process messages indefinitely. Then collect and return results for each processed message.

0 Answers
Related