I am trying to do stratified sampling and bootstraping afterwards. I have two data sets of nucleotides sequences (c,g,t,a) with two columns and many rows. dfu data has five columns and 37 rows with headings on the first row dfi data has five columns and more than thousands rows with headings on the first row.
I want to do random sampling from dfi data but the same as dfu data. so I tried to use dfu data as a map and group dfi according to the map.
mapper = dfu.set_index('Base')['Score'].apply(list).to_dict()
print(mapper)
j = dfi.groupby('Base').apply(lambda x: x.sample(n=mapper[x.name]))
print(j)
Mapper worked but from the grouping part, I get 'float' object is not iterable error
It would be great if I can get a help for this!
Thank you