How to get a list of all distinct prefixes in S3 bucket?

Viewed 4128

If I have a directory structure as below and the prefix is /folder1,

/folder1/folder11/folder12/folder13/*.files
               /folder21/folder22/folder23/*.files
               /folder31/folder32/*.files

I want to loop through these directories dynamically in order to read files in each of the leaf folder separately, i.e. I'd need a list

[
 /folder1/folder11/folder12/folder13/, 
 /folder1/folder21/folder22/folder23/,
 /folder1/folder31/folder32/
]

Is there a better way to get it other than loop through each prefix recursively, get next level prefix, concatenate, get next level, etc., until you get to the last (leaf) folder?

4 Answers

When listing objects from Amazon S3, if you specify Delimiter='/', then it will return a list of CommonPrefixes. This is effectively a list of subdirectories for the given Prefix.

However, I suggest that you do not think about directories. Instead, just loop through all objects and look at the Key to know the path of the object.

If you just want a list of paths that contain files, use this:

import boto3

BUCKET = 'my-bucket'

s3_resource = boto3.resource('s3')
folders = set()

# Find paths of all non-empty objects (to exclude zero-length 'folder' objects)
for object in s3_resource.Bucket(BUCKET).objects.all():
    if object.size > 0 and '/' in object.key:
        folders.add(object.key[:object.key.rfind('/')])

print (folders)

you can use this in a loop to get just the prefixes

def get_list_of_prefixes_from_prefix(bucket, prefix):
    """gets list of prefixes for given bucket and prefix"""
    list_of_prefixes = []
    paginator = boto3.resource('s3').meta.client.get_paginator('list_objects')
    for result in paginator.paginate(Bucket=bucket, Prefix=prefix, Delimiter='/'):
        # print(result)
        if 'CommonPrefixes' in result:
            prefixes = [f['Prefix'] for f in result['CommonPrefixes']]
            list_of_prefixes.extend(prefixes)
    return list_of_prefixes

list_of_prefixes = get_list_of_prefixes_from_prefix('my-bucket', 'my-prefix/')

Generate a S3 inventory report and and process the data using Athena. The Athena table structure for the report is also mentioned in the same AWS article. This is the serverless approach of doing the same.

import boto3

list_of_prefixes = []

def get_list_of_prefixes(bucket, prefix):
    global list_of_prefixes
    paginator = boto3.resource('s3').meta.client.get_paginator('list_objects')
    for result in paginator.paginate(Bucket=bucket, Prefix=prefix, Delimiter='/'):
        if 'CommonPrefixes' in result:
            for f in result['CommonPrefixes']:
                get_list_of_prefixes(bucket, f['Prefix'])
        else:
            list_of_prefixes.append(prefix)
    return list_of_prefixes

list_of_prefixes = get_list_of_prefixes('my-bucket', 'my-prefix/')

for i in list_of_prefixes:
    print(i)
Related