I am trying to achieve the following and although I think I am on the right track I am missing the last piece of the puzzle.
I am testing the performance of an algorithm and need to use combination of configurations/parameters. I am using Python 3.8 numpy/pandas/matplotlib as well as multiprocessing.
Every time the algorithm runs it produces a result. So having input config/params as I algorithm as A result as O I can say
- run 0 = I0 -> A -> O0
- run 1 = I1 -> A -> O1
- ...
- run n = In -> A -> On
All of the n runs can be executed in parallel with multiprocessing and I am very pleased of the CPU utilization at 100% and the speddup compared to the sequential execution but...
Say I can draw a chart for each one of the run and save each on of them in a separate figure, I would expect at the end of the process n figures, each one with the data related to a specific run. What happens instead is that for some reasons the plots overlap and so
- run 0 is ok
- run 1 has what run 1 should have plus run 0 data
- run 2 has what run 2 should have plus run 0 and run 1 data
(obviously not in order as they are running in parallel but hope you get the point)
I think this happens because to plot the charts I am using the generic object available in matplotlib
dataframe['run n'].plot(label=params_string, title="run 1").get_figure().savefig(fig_file)
Although that instruction runs in the separate context of each process for some reasons it overlaps in the charts. So at the end I have n charts but the get mixed up with the data of other execution.
I am pretty sure Pandas has a way to specify in what chart you want to plot the data but I can't just find the way.
Can anybody give a hand please?
Thank you, really appreciated in advance