I have two dataframes with sick and healthy people. They have different sizes and have different age distribution. I need to calculate some statistics, but I need to do it according to age-matched cohort. It means that from one dataset I need randomly select multiple times data with the same distribtuion of age as in initial.
Simple example. x array contains age of sick people, instead y array contains age of healthy people. I created two dataframes with index as id
import numpy as np
from numpy import random
x=random.randint(100, size=(1000))
y = random.randint(100, size=(700))
x_df = pd.DataFrame({'x':x})
y_df = pd.DataFrame({'y':y})
x_df['id'] = range(1, len(x_df) + 1)
x_df.set_index('id', inplace=True)
I want to randomly select multiple times the similar age distribution in two datasets. My script is below:
result = []
exclude_hlthy = []
for i in set(x_df['x']):
sick_ppl = x_df.index[x_df['x'] == i].tolist()
L_sick = len(sick_ppl)
if L_sick == 0:
continue
hlth_peers = y_df[y_df.y == i]
L_healthy = hlth_peers.shape[0]
if L_healthy < len(sick_ppl):
pass
else:
hlthy_subsample = list(np.random.choice([x for x in hlth_peers.index if not x in exclude_hlthy],
L_sick, replace = False))
exclude_hlthy += hlthy_subsample
result += hlthy_subsample
I want to get list of index from small dataset which will have same and similar (with step +1 +2....) age distribution. My script does not work. In real data it selects only two samples.