How to read parquet files from Azure Blobs into Pandas DataFrame?

Viewed 6058

I need to read .parquet files into a Pandas DataFrame in Python on my local machine without downloading the files. The parquet files are stored on Azure blobs with hierarchical directory structure. I am doing something like following and I am not sure how to proceed :

from azure.storage.blob import BlobServiceClient
blob_service_client = BlobServiceClient.from_connection_string(connection_string)

blob_client = blob_service_client.get_blob_client(container="abc", blob="/xyz/pqr/folder_with_parquet_files")

I have used dummy names here for privacy concerns. Assuming the directory "folder_with_parquet_files" contains 'n' no. of parquet files, how can I read them into a single Pandas DataFrame?

3 Answers

Hi you could use pandas and read parquet from stream. It colud be very helpful for small data set, sprak session is not required here. It could be the fastest way especially for testing purposes.

import pandas as pd
from io import BytesIO
from azure.storage.blob import ContainerClient

path = '/path_to_blob/..'
conn_string = <conn_string>
blob_name = f'{path}.parquet'

container = ContainerClient.from_connection_string(conn_str=conn_string, container_name=<name_of_container>)

blob_client = container.get_blob_client(blob=blob_name)
stream_downloader = blob_client.download_blob()
stream = BytesIO()
stream_downloader.readinto(stream)
processed_df = pd.read_parquet(stream, engine='pyarrow')

Here is a very similar solution, but slightly different using the new method azure.storage.blob._download.StorageStreamDownloader.readall:

from io import BytesIO
from azure.storage.blob import BlobServiceClient

blob_service_client = BlobServiceClient.from_connection_string(connection_string)
container_client = blob_service_client.get_container_client(container="parquet")

downloaded_blob = container_client.download_blob(upload_name)
bytes_io = BytesIO(downloaded_blob.readall())
df = pd.read_parquet(bytes_io)

print(df.head())

get_blob_to_bytes method can be used

Here the file is fetched from blob storage and held in memory. Pandas can then read this byte array as parquet format.

from azure.storage.blob import BlockBlobService
import pandas as pd
from io import BytesIO

#Source account and key
source_account_name = 'testdata'
source_account_key ='****************'

SOURCE_CONTAINER = 'my-data'
eachFile = 'test/2021/oct/myfile.parq'

source_block_blob_service = BlockBlobService(account_name=source_account_name, account_key=source_account_key)


f = source_block_blob_service.get_blob_to_bytes(SOURCE_CONTAINER, eachFile)
df = pd.read_parquet(BytesIO(f.content))
print(df.shape)
Related