I am training some PyTorch model on a Linux machine with 2 Tesla M60 GPUs and am using this example for splitting the training across the two GPUs using data parallelisation https://towardsdatascience.com/how-to-scale-training-on-multiple-gpus-dae1041f49d2
However, when the code gets to the point of calling the torch.multiprocessing.spawn function, it raises the
TypeError: cannot pickle '_thread.lock' object
I know it would be best practice to post a small reproducible example, however the codebase is very large, and I've been unable to reproduce this error on a small enough model to be posted here. The python error traceback also doesn't help much as it just refers to the internals of the multiprocessing library. I know this may sound a bit too general without a code sample, but I was just wondering if someone has encountered this same error when multiprocessing with PyTorch on a Linux machine, and if so, what are the things to look out for to try and fix it. Code runs absolutely fine on a single GPU, so it's proving very hard to debug.
If it helps, I'm using python 3.8.11 and torch 1.9.0, CUDA version 11.1