I have this set:
df=pd.DataFrame({'user':[1,1,2,2,2,3,3,3,3,3,4,4],
'date':['1995-09-01','1995-09-02','1995-10-03','1995-10-04','1995-10-05','1995-11-07','1995-11-08','1995-11-09','1995-11-10','1995-11-15','1995-12-18','1995-12-20'],
'type':['a','a','b','a','c','a','b','a','b','b','a','b']})
Which gives me:
user date type
1 1995-09-01 a
1 1995-09-02 a
2 1995-10-03 b
2 1995-10-04 a
2 1995-10-05 c
3 1995-11-07 a
3 1995-11-08 b
3 1995-11-09 a
3 1995-11-10 b
3 1995-11-15 b
4 1995-12-18 a
4 1995-12-20 b
I want to create a new column where the count of a values on "type" column is shown, grouped by column "user""
Here is the expected outcome:
user date type cta_a
1 1995-09-01 a 2
1 1995-09-02 a 2
2 1995-10-03 b 1
2 1995-10-04 a 1
2 1995-10-05 c 1
3 1995-11-07 a 2
3 1995-11-08 b 2
3 1995-11-09 a 2
3 1995-11-10 b 2
3 1995-11-15 b 2
4 1995-12-18 a 1
4 1995-12-20 b 1
I tried the following but it did not work.
df['ct_a'] = df.groupby('user')[df['type']== 'a'].transform('count')