I had the following function in pandas 0.17:
df['numberrows'] = df.groupby(['column1','column2','column3'], as_index=False)[['column1']].transform('count').astype('int')
But I upgraded pandas today and now I get the error:
File "/usr/local/lib/python3.4/dist-packages/pandas/core/internals.py",line 3810, in insert raise ValueError('cannot insert {}, already exists'.format(item))
ValueError: cannot insert column1, already exists
What has changed in the update which causes this function to not work anymore?
I want to groupby the columns and add a column which has the amount or rows of the groupby.
If what I did before was not a good function, another way of grouping while getting the amount of rows that were grouped is also welcome.
EDIT:
small dataset:
column1 column2 column3
0 test car1 1
1 test2 car5 2
2 test car1 1
3 test4 car2 1
4 test2 car1 1
outcome would be:
column1 column2 column3 numberrows
0 test car1 1 2
1 test2 car5 2 1
3 test4 car2 1 1
4 test2 car1 1 1