How to select data with the same distribution from another dataframe

Viewed 138

I have two dataframes with sick and healthy people. They have different sizes and have different age distribution. I need to calculate some statistics, but I need to do it according to age-matched cohort. It means that from one dataset I need randomly select multiple times data with the same distribtuion of age as in initial.

Simple example. x array contains age of sick people, instead y array contains age of healthy people. I created two dataframes with index as id

import numpy as np
from numpy import random
x=random.randint(100, size=(1000))
y = random.randint(100, size=(700))
x_df = pd.DataFrame({'x':x})
y_df = pd.DataFrame({'y':y})
x_df['id'] = range(1, len(x_df) + 1)
x_df.set_index('id', inplace=True)

I want to randomly select multiple times the similar age distribution in two datasets. My script is below:

result = []
exclude_hlthy = []
for i in set(x_df['x']):
    sick_ppl = x_df.index[x_df['x'] == i].tolist()
    L_sick = len(sick_ppl)
    if L_sick == 0:
        continue
    hlth_peers = y_df[y_df.y == i]
    L_healthy = hlth_peers.shape[0]
    if L_healthy < len(sick_ppl):
        pass
    else:
        hlthy_subsample = list(np.random.choice([x for x in hlth_peers.index if not x in exclude_hlthy], 
                                                    L_sick, replace = False))
        exclude_hlthy += hlthy_subsample
        result += hlthy_subsample

I want to get list of index from small dataset which will have same and similar (with step +1 +2....) age distribution. My script does not work. In real data it selects only two samples.

0 Answers
Related