I have a dataframe with the users' history in online shop. Example:
In [1]: a = pd.DataFrame([[1, 'view', 'a'], [1, 'cart', 'b'], [2, 'cart','b'], [2, 'cart','c'], [2, 'view','d'],
[2, 'purchase','d'], [2, 'view','e'], [2, 'cart','e']],
columns=['user_session', 'event_type', 'product_id'])
In [2]: df
Out[2]:
user_session event_type product_id
0 1 view a
1 1 cart b
2 2 cart b
3 2 cart c
4 2 view d
5 2 purchase d
6 2 view e
7 2 cart e
There can be more purchases pro one user_session. I need to delete ALL further rows in a session as soon as first purchase occurs. The partial solution I found here: Removing rows after a certain string in pandas and it is:
df.loc[:(df['event_type'] == 'purchase').idxmax()]
But I need to iterate thru a huge dataset with millions of rows. Is it a good idea to use for loop here?It should be probably better opportunity.
Another way would be probably to build up a list of indexes of rows that I want to delete as mentioned here: dropping a row while iterating through pandas dataframe
for i in df.index:
....
if {make your decision here}:
indexes_to_drop.append(i)
....
df.drop(df.index[indexes_to_drop], inplace=True )
But again, is there any other way?
Many thanks in advance!