Use checksum to verify integrity of uploaded and downloaded files from AWS S3 via Django

Viewed 538

Using Django I am trying to upload multiple files in AWS S3. The file size may vary from 500 MB to 2 GB. I need to check the integrity of both uploaded and downloaded files.

I have seen that using PUT operation I can upload a single object up to 5GB in size. They also provide 'ContentMD5' option to verify the file. My questions are:

  • Should I use PUT option if I upload a file larger than 1 GB? Because generating MD5 checksum of such file may exceed system memory. How can I solve this issue? Or is there any better solution available for this task?

To download the file with checksum, AWS has get_object()function. My questions are:

  • Is it okay to use this function to downlaod multiple files?
  • How can I use it to download multiple files with checksum from S3? I look for some example but there is noting much. Now I am using following code to download and serve multiple files as a zip in Django. I want to serve the files as zip with checksum to validate the transfer.
    s3 = boto3.resource('s3', aws_access_key_id=base.AWS_ACCESS_KEY_ID, aws_secret_access_key=base.AWS_SECRET_ACCESS_KEY)
    bucket = s3.Bucket(base.AWS_STORAGE_BUCKET_NAME)
    s3_file_path = bucket.objects.filter(Prefix='media/{}/'.format(url.split('/')[-1]))

    # set up zip folder
    zip_subdir = url.split('/')[-1]
    zip_filename = zip_subdir + ".zip"
    byte_stream = BytesIO()
    zf = ZipFile(byte_stream, "w")

    for path in s3_file_path:
        s3_url = f"https://%s.s3.%s.amazonaws.com/%s" % (base.AWS_STORAGE_BUCKET_NAME,base.AWS_S3_REGION_NAME,path.key)
        file_response = requests.get(s3_url)
        if file_response.status_code == 200:
            try:
                tmp = tempfile.NamedTemporaryFile()
                print(tmp.name)
                tmp.name = path.key.split('/')[-1]
                f1 = open(tmp.name, 'wb')
                f1.write(file_response.content)
                f1.close()
                zip_path = os.path.join('/'.join(path.key.split('/')[1:-1]), tmp.name)
                zf.write(tmp.name,zip_path)
            finally:
                os.remove(tmp.name)
    zf.close()
    response = HttpResponse(byte_stream.getvalue(), content_type="application/x-zip-compressed")
    response['Content-Disposition'] = 'attachment; filename=%s' % zip_filename

I am learning about AWS S3 and this is the first time I am using it. I would appreciate any kind of suggestion regarding this problem.

0 Answers
Related