Speed difference reading parquet files between pandas and pyarrow

Viewed 1720

I am trying to understand why there is a such a difference in speed between reading a parquet file directly to Pandas using pd.read_parquet and Pyarrow.parquet pq.read_table() It seems strange as I believe Pandas is using Pyarrow under the hood.

What am I missing?

pandas version 1.1.5

pyarrow version 3.0.0

For example reading the same file from S3

start = datetime.datetime.now()
pd_df = pd.read_parquet("s3://" + bucket_name + "/" + key)
total = datetime.datetime.now() - start
print(total)
0:00:27.634829

and

start = datetime.datetime.now()
pa_tab = pq.read_table("s3://" + bucket_name + "/" + key)
total = datetime.datetime.now() - start
print(total)
0:16:18.545071

Edit

I tried reading the file from the local drive and get much better performance from pyarrow.

start = datetime.datetime.now()
pd_df = pd.read_parquet("~/injury_problem/athlete_18931.parquet")
total = datetime.datetime.now() - start
print(total)
0:00:27.224333

vs

start = datetime.datetime.now()
pa_tab = pq.read_table("/Users/warwick/injury_problem/athlete_18931.parquet")
total = datetime.datetime.now() - start
print(total)
0:00:19.376693
0 Answers
Related