What are the best practices for training one neural net on more than one GPU on one machine?
I'm a little confused by the different options from nn.DataParallel vs putting different layers on different GPUs with .to('cuda:0') and .to('cuda:1'). I see in the Pytorch docs the latter method the date was 2017. Is there a standard or does it depend on preference or the type of model?
Method 1
class ToyModel(nn.Module):
def __init__(self):
super(ToyModel, self).__init__()
self.net1 = torch.nn.Linear(10, 10)
self.relu = torch.nn.ReLU()
self.net2 = torch.nn.Linear(10, 5)
def forward(self, x):
x = self.relu(self.net1(x))
return self.net2(x)
model = ToyModel().to('cuda')
model = nn.DataParallel(model)
Method 2
class ToyModel(nn.Module):
def __init__(self):
super(ToyModel, self).__init__()
self.net1 = torch.nn.Linear(10, 10).to('cuda:0')
self.relu = torch.nn.ReLU()
self.net2 = torch.nn.Linear(10, 5).to('cuda:1')
def forward(self, x):
x = self.relu(self.net1(x.to('cuda:0')))
return self.net2(x.to('cuda:1'))
I'm not sure there aren't more ways Pytorch provides to train on more than one GPU. Both of these methods seem to cause my system to freeze depending on what model I use them. In Jupyter the cell stays at a [*] and if I don't restart the kernel the screen freezes and I have to do a hard reset. A few tutorials on multi-gpu cause my system to hang and freeze like this.