How to fix "cannot pickle '_thread.lock' object" error when using torch.multiprocessing.spawn

Viewed 709

I am training some PyTorch model on a Linux machine with 2 Tesla M60 GPUs and am using this example for splitting the training across the two GPUs using data parallelisation https://towardsdatascience.com/how-to-scale-training-on-multiple-gpus-dae1041f49d2

However, when the code gets to the point of calling the torch.multiprocessing.spawn function, it raises the

TypeError: cannot pickle '_thread.lock' object

I know it would be best practice to post a small reproducible example, however the codebase is very large, and I've been unable to reproduce this error on a small enough model to be posted here. The python error traceback also doesn't help much as it just refers to the internals of the multiprocessing library. I know this may sound a bit too general without a code sample, but I was just wondering if someone has encountered this same error when multiprocessing with PyTorch on a Linux machine, and if so, what are the things to look out for to try and fix it. Code runs absolutely fine on a single GPU, so it's proving very hard to debug.

If it helps, I'm using python 3.8.11 and torch 1.9.0, CUDA version 11.1

0 Answers
Related