Pytorch: Load a transformer model directly from torch checkpoint without loading pre-initialized weights

Viewed 497

I have a memory constraint while loading a model from a torch checkpoint for inference: What I have is the following:

A torch model (xlm-roberta)
and a checkpoint (xlm-roberta-checkpoint.pth)

What normally happens is that we load the xlm-roberta as follows:

model = AutoModel.from_pretrained('xlm-roberta-base')
checkpoint = torch.load(PATH)
model.load_state_dict(checkpoint['model_state_dict'])

But the problem arises when loading the checkpoint; the pre-trained model itself is quite large, so both the checkpoint, and the model cannot fit in the memory and the process dies out.
Is there a way to directly load the checkpoint into the model class without initializing some random/pre-trained weights beforehand?

0 Answers
Related