So I have an input that consists in a dataset and several ML algorithms (with parameter tuning) using scikit-learn. I have tried quite a few attempts on how to execute this as efficiently as possible but at this very moment I still don't have the proper infrastructure to assess my results. However, I lack some background on this area and I need help to get things cleared up.
Basically I want to know how the tasks are distributed in a way that exploits as much as possible all the available resources, and what is actually done implicitly (for instance by Spark) and what isn't.
I need to train many different Decision Tree models (as many as the combination of all possible parameters), many different Random Forest models, and so on...
In one of my approaches, I have a list and each of its elements corresponds to one ML algorithm and its list of parameters.
spark.parallelize(algorithms).map(lambda algorihtm: run_experiment(dataframe, algorithm))
In this function run_experiment I create a GridSearchCV for the corresponding ML algorithm with its parameter grid. I also set n_jobs=-1 in order to (try to) achieve maximum parallelism.
In this context, on my Spark cluster with a few nodes, does it make sense that the execution would look somewhat like this?
Or there can be one Decision Tree model and also one Random Forest model running in the same node? This is my first experience using a cluster environment so I am a bit confused on how to expect things to work.
On the other hand, what exactly changes in terms of execution, if instead of the first approach with parallelize, I use a for loop to sequentially iterate through my list of algorithms and create the GridSearchCV using databricks's spark-sklearn integration between Spark and scikit-learn? The way it's illustrated in the documentation it seems something like this:
Finally, with regards to this second approach, using the same ML algorithms but instead with Spark MLlib instead of scikit-learn, would the whole parallelization/distribution be taken care of?
Sorry if most of this is a bit naive, but I really appreciate any answers or insights on this. I wanted to understand the basics before actually testing in the cluster and playing with task scheduling parameters.
I am not sure whether this question is more suitable here or on CS stackexchange.


