Fastest way to sample Pandas Dataframe?

Viewed 4211

First, I want to take random samples from three dataframes (150 rows each) and concat the results. Second, I want to repeat this process as many times as possible.

For part 1 I use the following function:

def get_sample(n_A, n_B, n_C):
    A = df_A.sample(n = n_A, replace=False)
    B = df_B.sample(n = n_B, replace=False)
    C = df_C.sample(n = n_C, replace=False)
    return pd.concat([A, B, C])

For part 2 I use the following line:

results = [get_sample(5,5,3) for i in range(n)] 

Currently with n = 50.000 the analysis takes about 1 minute and 40 seconds on my MacBook. Any advise on how to improve the speed of this process is welcome!

PM the three dataframes (df_A, df_B, df_C) differ only in one categorical feature. The challenge is that I want a specific number samples from each category.

2 Answers

Working with numpy ndarrays should be faster, since pandas itself is built on numpy. The sampling can be done with: numpy.random.choice, as explained here. That should work as an equivalent of pd.sample. Then you can switch back from numpy to pandas.

In your case it should pay off to work with numpy arrays instead of pandas dataframes (as noted already by Leevo).

Numpy arrays are simpler objects than pandas dataframes (the absence of row/column labels in numpy arrays is a prime example). As a result numpy arrays allow operations such as concatenation to be performed faster. The time difference is usually negligible when you're performing just a few concatenations within a larger script. However in your case where you're doing concatenations within a many-iterations loop, time differences can accumulate and become significant.

Try the following:

import pandas as pd
import numpy as np

# Initialize example dataframes
df_A = pd.DataFrame(np.random.rand(150, 10))
df_B = pd.DataFrame(np.random.rand(150, 10))
df_C = pd.DataFrame(np.random.rand(150, 10))

# Initialize constants
n_A = 5
n_B = 5
n_C = 3
n = 10000

# Reduce dataframes to numpy arrays
arr_A = df_A.values
arr_B = df_B.values
arr_C = df_C.values

# Perform sampling on numpy arrays
def get_sample():
    A = arr_A[np.random.choice(arr_A.shape[0], n_A, replace=False)]
    B = arr_B[np.random.choice(arr_B.shape[0], n_B, replace=False)]
    C = arr_C[np.random.choice(arr_C.shape[0], n_C, replace=False)]
    return np.concatenate([A, B, C])
results = [get_sample() for i in range(n)]
Related