Loading files from s3 with subfolders with python

Viewed 3618

I'm trying to load a csv file in pandas from a s3 bucket in aws. Boto3 seems to fall short in providing functionalities for loading files from subfolders. Let's say i have the following path in s3: bucket1/bucketwithfiles1/file1.csv

How do i specify how to load file1.csv? I know s3 doesn't have a directory structure.

import boto3
import pandas as pd

s3 = boto3.client('s3')
obj = s3.get_object(Bucket='/bucket1/creditdefault-ff.csv')

df = pd.read_csv(obj['Body'])
2 Answers

You may have multiple files in a Bucket, each one is identified by a Key (which is the path to the file in S3). So, you want to get a dataframe for all the files (all the keys) in a single Bucket.

s3 = boto3.client('s3')
obj = s3.get_object(Bucket='my-bucket', Key='my-file-path')

df = pd.read_csv(obj['Body'])

In the case you have multiple files, you'll need to combine boto3 methods named list_object_v2 (to get keys in the bucket you specified), and get_object using a loop on retrieved keys to get all your files.

Then, it might be useful to use the Prefix parameter of the list_object_v2 method to filter on a subfolder in your bucket.

It's a bit of code to write each time you need it, so you can find small Python module to do it for you and get extra features like the pandas_aws python package:

from pandas_aws import get_client
from pandas_aws.s3 import get_df_from_keys

s3 = get_client('s3') # you can use your boto3 s3 client if you already hase instanciated one

df = get_df_from_keys(s3, "my-bucket", "my-subfolder/", suffix='.csv')

Actually, it calls the same boto3 methods but filters results based on the provided suffix. The Dataframe construction is also embedded and based on the storage format.

See the package here: https://github.com/FlorentPajot/pandas-aws

Related