PyGitHub - Error 403 {"message": "This API returns blobs up to 1 MB in size

Viewed 1338

I am working on a Python script to automatically create files on Github and if they exist, update them. I am using the module PyGithub with the logic below.

The problem I have is that when I try to update a file that is bigger that 1Mb, I get:

github.GithubException.GithubException: 403 {"message": "This API returns blobs up to 1 MB in size. The requested blob is too large to fetch via the API, but you can use the Git Data API to request blobs up to 100 MB in size.", "errors": [{"resource": "Blob", "field": "data", "code": "too_large"}], "documentation_url": "https://developer.github.com/enterprise/2.20/v3/repos/contents/#get-contents"}

I've tried deleting the files and recreating them but I understand that just the fact of reading the file triggers the error. I've tried several option but nothing works.

I am stuck. Thanks for the help

try:
    repo.create_file(file_path, "elastic_backups", bk_object.text, branch="master")
    print('creating new file ',file_path)
except:
    contents = repo.get_contents(file_path, ref="master")
    repo.update_file(contents.path, "updated elastic backup", bk_object.text, contents.sha, branch="master")
    print(file_path, ' UPDATED')
       
1 Answers

I saw a nice snippet of code that may solve your issue. At least it worked for me. Taken from this GitHub issue.

As many people in other questions similar to this one answered: for files larger than 1MB you have to you use the GitHub Data API, which stores those files as blobs (binary large objects), following the documentation you can learn more about that. For every blob you get a SHA1 associate to it, which is computed and then stored in the blob object. So if you want to retrieve this larger file you would need to provide the SHA1 to retrieve the blob from GitHub. The snippet below does exactly that:

Define a function called get_blob_content() which is going to implement all the logic I mentioned above:

def get_blob_content(repo, branch, path_name):
    # first get the branch reference
    ref = repo.get_git_ref(f'heads/{branch}')
    # then get the tree
    tree = repo.get_git_tree(ref.object.sha, recursive='/' in path_name).tree
    # look for path in tree
    sha = [x.sha for x in tree if x.path == path_name]
    if not sha:
        # well, not found..
        return None
    # we have sha
    return repo.get_git_blob(sha[0])

On your Except block, call the get_blob_content().

Related