I want to transform each group in a pandas' DataFrame. By group I mean not a single column of the DataFrame, but the entire group. Here is an example of what I mean:
df = pd.DataFrame({'A' : ['foo', 'bar', 'foo', 'bar', 'foo', 'bar'],
'B' : ['one', 'one', 'two', 'two', 'one', 'two'],
'C' : [ 1, 5, 5, 2, 6, 5],
'D' : [ 2.0, 5., 8., 1., 2., 9.]})
def transformation_function(group: pd.DataFrame) -> pd.DataFrame:
group = group.copy()
if all(group.B == 'one'):
group.D[group.C>2] = group.D[group.C>2] + 1
else:
group.A = 'new'
return group
df.groupby('B').transform(transformation_function)
where I would expect
pd.DataFrame({'A' : ['foo', 'bar', 'new', 'new', 'foo', 'new'],
'B' : ['one', 'one', 'two', 'two', 'one', 'two'],
'C' : [ 1, 5, 5, 2, 5, 5],
'D' : [ 2.0, 6., 8., 1., 3., 9.]})
as a result. Now, I get the
AttributeError: 'Series' object has no attribute 'B'
which does not make sense to me, because the documentation explicitly states
Call function producing a like-indexed DataFrame on each group and return a DataFrame having the same indexes as the original object filled with the transformed values
I am aware the all examples work Series-based, like df.groupby('B')['a_column_name'].transform(change_fct), but something like that is not possible if one needs all columns for the transformation function.
So, how do I get what I expect, using pandas' method chaining?