Pandas: First column access of dataframe very slow

Viewed 46

I have merged two large dataframes and am consecutively accessing pairs of columns (as in df[[col1, col2]]). First time, the retrieval is very slow while the following ones are as fast as "normal". Why is this and how can I avoid it?

Ex.

from time import time
import numpy as np
import pandas as pd

# generate two big dataframes with common column values
df1 = pd.DataFrame(np.random.rand(800000, 3))
df2 = df1.sample(len(df1)).reset_index(drop=True)
# merge on column 0
df = df1.merge(df2, on=0)

# time retrieval of suffixed columns
start = time()
df[['1_x', '1_y']];
print('1st access:', time() - start)

start = time()
df[['2_x', '2_y']];
print('2nd access:', time() - start)

The output is

1st access: 2.8012924194335938
2nd access: 0.003415822982788086
0 Answers
Related