I am trying to understand why there is a such a difference in speed between reading a parquet file directly to Pandas using pd.read_parquet and Pyarrow.parquet pq.read_table() It seems strange as I believe Pandas is using Pyarrow under the hood.
What am I missing?
pandas version 1.1.5
pyarrow version 3.0.0
For example reading the same file from S3
start = datetime.datetime.now()
pd_df = pd.read_parquet("s3://" + bucket_name + "/" + key)
total = datetime.datetime.now() - start
print(total)
0:00:27.634829
and
start = datetime.datetime.now()
pa_tab = pq.read_table("s3://" + bucket_name + "/" + key)
total = datetime.datetime.now() - start
print(total)
0:16:18.545071
Edit
I tried reading the file from the local drive and get much better performance from pyarrow.
start = datetime.datetime.now()
pd_df = pd.read_parquet("~/injury_problem/athlete_18931.parquet")
total = datetime.datetime.now() - start
print(total)
0:00:27.224333
vs
start = datetime.datetime.now()
pa_tab = pq.read_table("/Users/warwick/injury_problem/athlete_18931.parquet")
total = datetime.datetime.now() - start
print(total)
0:00:19.376693