I would like to aggregate a dataframe using a kind of intersection that goes like this. One column of the dataframe is made up of lists of varying length. Using a threshold, if two value are close enough, I'd like the mean mean to be kept in the new dataframe. So that this :
entry = pd.DataFrame({'name' : ['a', 'b', 'a', 'b'], 'values' : [[1.01,2.4,3.7], [1.01,2.4,3.7], [1.0,2.5,3.73], [1.01,3.7, 5., 6.]]})
when grouped by name and aggregated using a threshold of 0.05, yelds this :
result = pd.DataFrame({'name' : ['a', 'b'], 'values' : [[1.005,3.715], [1.01,3.7]]})
I tried using these but it doesn't work :
def inters(x, y, tolerance = 0.05):
res = []
for x_i in x:
for y_i in y:
if(abs(x_i-y_i)<tolerance):
res.append(x_i)
return res
def fct(series):
return reduce(lambda x, y : inters(x,y), series)
alldf.groupby(['dog', 'input', 'nwin']).agg({'sig_pulses' : ['inters', fct]})