Does torch.distributed support point-to-point communication for GPU?

Viewed 177

I am looking into how to do point-to-point communication with multiple GPUs on separate nodes in PyTorch.

As of version 1.10.0, the documentation page for PyTorch says question marks for send and recv for GPU with the MPI backend. What does this mean? If anyone has successfully set up PyTorch so that torch.distributed allows point-to-point communication on multiple GPUs, please let me know and how you set it up. Specifically, which MPI are you using? What about the versions of pyTorch and Cuda?

1 Answers

I guess I'll post what I have learned so far.

Pytorch does seem to support p-to-p communication with MPI on GPU. However, this requires you to have a Cuda-aware MPI. (If your MPI isn't Cuda-aware, you'll need to build MPI from source with a specific parameter). In addition, if your Pytorch doesn't have MPI enabled, you need to compile Pytorch from source with MPI installed. This seems a very complicated route to go.

However, it seems the documentation I linked to is misleading. Looking at the release note, Pytorch supports send/recv in NCCL backend since 1.8.0... That being said, I have tried doing send/recv with NCCL but it throws errors saying NCCL are getting invalid arguments. I'm not sure if it's my problem or there are still bugs in pytorch distributed code.

Related