Unable to load weights from pytorch checkpoint after splitting pytorch_model.bin into chunks

Viewed 1563

I need to transfer a pytorch_model.bin of a pretrained deeppavlov ruBERT model but I have a file size limit. So I split it into chunks using python, transferred and reassembled in the correct order. However, the size of the file increased, and when I tried to load the resulting file using BertModel.from_pretrained(pytorch_model.bin) I received an error:

During handling of the above exception, another exception occurred:
OSError: Unable to load weights from pytorch checkpoint <...>

So my question is: is it actually possible to split the file like that? I could possibly have a mistake in the way I split and reassemble the file. However, this could also be some version mismatch.

My python code to get chunks:

chunk_size = 40000000
file_num = 1
with open("pytorch_model.bin", "rb") as f:
    chunk = f.read(chunk_size)
    while chunk:
        with open("chunk_" + str(file_num), "wb") as chunk_file:
            chunk_file.write(chunk)
        file_num += 1
        chunk = f.read(chunk_size)

Code to reassemble one file:

chunks = !ls | grep chunk_
chunks = sorted(chunks, key=lambda x: int(x.split("_")[-1]))

for chunk in chunks:
    with open(chunk, "rb") as f:
        contents = f.read()
    if chunk == chunks[0]:
        write_mode = "wb"
    else:
        write_mode = "ab"
    with open("pytorch_model.bin", write_mode) as f:
        f.write(contents)

python 3.7.0, torch 1.5.1, transformers 4.2.2. I have no way to move files bigger than 40 MB.

TIA for your help!

3 Answers

Those who are new to this issue I just figured it out and save your time

What is this error about? ==> When you run the model for the first time it downloads some files { pytorch_model.bin } and if your internet is broken accidentally between processes it will continue running the pipeline file without completely downloading that pytorch_model.bin file so it will raise this issue.

Steps : 
1 ] Go to C:// Users / UserName / .cache
2 ] Delete .cache folder
3 ] And Done Just Run The Model Once Again......

I suggest leaving python out of this.

Use command line split and cat to split the large file and to join the splits into a single file on the other side (this thread shows how).

I suggest you use md5sum (or other checksum functions) to verify that the pytorch_model.bin file you assembled at the receiving side is indeed identical to the original one.

I checked with my team about the versions of transformers and pytorch used when the model was saved. It was different from the versions I was using to load the model. So I installed the versions used when the model was saved, and then re-tried the loading. It worked.

Related