I am trying to train a model (convLSTM) using Pytorch for video prediction. My "videos" are a list of images of dimensions 350x350 (frames), and I give N frames as input and need to predict N frames in output.
For train my network I am currently using a NVIDIA V100 with 12 cores and 12GB of RAM. My inputs and ouputs are 5-th order tensors of size [batch_size, n_frames, channels, height, width].
The problem is that even setting the bare minimum, i.e.: num_workers=1,batch_size = 1, size-training-set= 1 (just 1 video sample of 16 frames! - note that with lower than 16 frames works) I get this error:
RuntimeError: CUDA out of memory. Tried to allocate 24.00 MiB (GPU 0; 32.00 GiB total capacity; 28.14 GiB already allocated; 96.54 MiB free; 28.55 GiB reserved in total by PyTorch) If reserved memory is >> allocated memory try setting max_split_size_mb to avoid fragmentation.
I really don't know how to get rid of this error. My training size is about more than 10K samples of 20 frames each and I cannot even run my code with 1 sample.
I saw that this error is common, but cannot find a way to solve it (since everything is already at the minimum..). I tried to check my memory usage with nvidia-smi and torch.cuda.memory_summary() but could not get any insight. Also, torch.cuda.empty_cache() does not help.
I see however that pytorch allocate a lot of space for itself, is it normal? In the code I don't do nothing weird, I just check if cuda is available as device and then set
model.to(device)
if multi_gpu:
model = nn.DataParallel(model)
the error generates at this point in my training loop:
prediction = model(inputs,
input_frames = train_data.n_frames_input,
future_frames = train_data.n_frames_output,
output_frames = train_data.n_frames_output,
teacher_forcing = True,
scheduled_sampling_ratio = scheduled_sampling_ratio)
Any insights?
EDIT: resizing the pictures to lower dimensions (half of the original quality) seems to help.