I have a df with columns that represents a stratum (strat). I want to loop over those stratum and pull out rows to a new df, df_sample. I want to pull out all rows in a stratum if cases are few.
I've tried the below, and it works. But I wonder if there is a better solution to this problem. Perhaps pd.concat is slow when I later use the real much larger data for example.
df=pd.DataFrame({'ID': range(0,120),
'strat': ['A', 'B', 'B', 'A', 'B', 'A', 'D', 'A', 'B', 'C',
'A', 'D', 'A', 'A', 'A', 'D', 'F', 'D', 'F', 'C',
'B', 'A', 'A', 'C', 'A', 'A', 'B', 'D', 'B', 'C',
'C', 'A', 'C', 'A', 'C', 'A', 'D', 'C', 'C', 'A',
'B', 'F', 'F', 'C', 'B', 'D', 'A', 'A', 'B', 'B',
'A', 'C', 'A', 'A', 'F', 'A', 'A', 'B', 'A', 'D',
'C', 'B', 'B', 'A', 'B', 'C', 'B', 'A', 'D', 'B',
'B', 'A', 'A', 'C', 'D', 'F', 'F', 'A', 'B', 'C',
'F', 'B', 'D', 'A', 'A', 'F', 'B', 'D', 'B', 'A',
'F', 'D', 'A', 'A', 'C', 'B', 'B', 'C', 'C', 'B',
'F', 'A', 'A', 'B', 'B', 'B', 'F', 'A', 'B', 'C',
'A', 'A', 'A', 'B', 'B', 'A', 'A', 'A', 'B', 'B']})
df_sample=pd.DataFrame()
for i in df.strat.unique():
temp=df[df['strat']==i]
if len(temp) < 21:
strat=temp.sample(len(temp))
elif len(temp) > 20:
strat = temp.sample(frac=0.5)
df_sample=pd.concat([df_sample, strat])