Best described with an example
import pandas as pd
df = pd.DataFrame({
'a' : ['A','B','C','A','B','C','A','B','C'],
'b': [1,2,3,4,5,6,7,8,9]}
)
And i want to create a column that contains in a list the elements of column b by group of column a
resulting in the following
a b c
0 A 1 [1, 4, 7]
1 A 4 [1, 4, 7]
2 A 7 [1, 4, 7]
3 B 2 [2, 5, 8]
4 B 5 [2, 5, 8]
5 B 8 [2, 5, 8]
6 C 3 [3, 6, 9]
7 C 6 [3, 6, 9]
8 C 9 [3, 6, 9]
I can do this with groupby and apply or agg and then joining the dataframes like so
df_tmp = df.groupby('a')['b'].agg(list).reset_index()
df.merge(df_tmp, on='a')
But i would also be expecting to do the same with transform
df['c'] = df.groupby('a')['b'].transform(list)
but the column c is the same as column b
Also the following
df.groupby('a')['b'].transform(lambda x: len(x))
return a series with the values 3 i.e. the length of the grouped elements is 3 (to be expected)
Also this
df.groupby('a')['b'].transform(lambda x: list(x))
does not provide the expected result.
So to my question, how can i obtain the desired result with groupby and tranform
pandas version is 1.0.5