Unfortunately, I am not able to share a MWE because with simple examples I cannot reproduce what happens when my full code runs on larger files.
The workflow can be so summarized:
- each file in a folder is processed by a separate process
P_imanaged by amultiprocessing.Manager()instance; - inside of
P_iascikit.model_selection.RandomizedSearchCVis run with three parallel threads managed byjoblib - when a process
P_iterminates with success, I get the following warning:
/home/user/miniconda3/envs/e/lib/python3.8/site-packages/joblib/externals/loky/backend/resource_tracker.py:318:
UserWarning: resource_tracker: There appear to be 60 leaked folder objects to clean up at shutdown
I could not find anything explaining that warning and how to avoid it. Is there anyone who happens to know what the warning is about?
"pseudo code" follows.
[...]
manager = multiprocessing.Manager()
return_dict = manager.dict()
for i, subf in enumerate(subfolders):
par_jobs = [] # list of spawned processes
filenames = [fn for fn in os.listdir( join(in_fld, subf)) if fn.endswith(".pickle")]
folder_results = [] # list of dataframes for each file in current folder
n_spawned = 0
for j, fn in enumerate(filenames):
abs_fn = join(in_fld, subf, fn)
p = multiprocessing.Process(target=process_file, args=(abs_fn, subf, labels, config, return_dict))
par_jobs.append(p)
p.start()
# <
for proc in par_jobs:
proc.join()
[...]
In process_file there is a:
[...]
grid = RandomizedSearchCV(pipe, parameters, n_iter=100, ..., n_jobs=3)
grid.fit()
[...]
During the execution there are no errors, only warnings:
ConvergenceWarning: Liblinear failed to converge, increase the number of iterations.
warnings.warn(
There are many questions and posts related to leaked semaphores. That problem seems to be related to RAM issues. I monitored the RAM usage of my application and it stayed well below the available RAM (around 7GB out of 32GB available).
There are 100GB left on disk, the 12 files (dataframes) processed by the app are around 70MB each.