My question is relevant to my previous one pytorch allocate memory for small size tensor on cpu and gpu but got error on a node with more than 400 GB. But, it is different so I created a new thread.
In this question, I have changed size of the input tensor size.
import torch
from torch import nn
import numpy as np
num_embedding, num_dim = 14000, 300
embedding = nn.Embedding(num_embedding, num_dim)
row, col = 8000, 302
t = [[x for x in range(col)] for _ in range(row)]
t1 = torch.tensor(t)
print(t1.shape) # torch.Size([8000, 302])
type(t1), t1.device, (t1.nelement() * t1.element_size())/(1024**3) # (torch.Tensor, device(type='cpu'), 0.01800060272216797)
tt = embedding(t1)
embedding.forward(t1)
t2 = t1.cuda()
t2.device, t2.shape, t2.grad, t2.nelement(), t2.element_size(), (t2.nelement() * t2.element_size())/(1024**3) # (device(type='cuda', index=0), torch.Size([8000, 302]), None, 2416000, 8, 0.01800060272216797)
embedding_cuda = embedding.cuda()
torch.cuda.empty_cache()
embedding_cuda(t2) # RuntimeError: CUDA out of memory. Tried to allocate 2.70 GiB (GPU 0; 11.17 GiB total capacity; 7.19 GiB already allocated; 2.01 GiB free; 8.88 GiB reserved in total by PyTorch)
Why the small size tensor (0.018 GB) can be allocated to cpu but cannot be allocated to gpu on the same node (p2.8xlarge) ? why it requires 2.7 GB, which 100 times larger than its original size at least ?
I have checked most posts at https://stackoverflow.com/search?q=RuntimeError%3A+CUDA+out+of+memory.+Tried+to+allocate+GiB but, none of them can help me about this.