How to calculate total number of seconds a detected class appear in frame with pandas?

Viewed 58

I am working on object detection project where my task is calculate for exactly how many seconds particular class was in the frame. I have a csv file of detected classes with their timestamp that looks like this:

enter image description here

I can input this csv into a pandas dataframe to calculate their timestamp range as finaltimestamp-intialtimestamp. But the catch is here is: suppose one class, let say HP, made an appearance for 5 seconds. After that, a new class kellogs is introduced and then HP reenters the frame.

Following the above final-intial logic fails here as there is a time gap after the same class appears again. How to deal with this in pandas? I'm aware of .groupby() and .valueCounts() but they can't solve this problem directly.

Example data:

           cat           time    
0           HP       06:35:03
1           HP       06:35:04
2         kellogs    06:35:42
3         kellogs    06:35:43
4           HP       06:35:45

Expected output

          cat       time
0         HP      00:00:03
1       kellogs   00:00:02

The output above should return as much time that each class was present in the frame. So in the above example, HP has 3 seconds and kellogs 2 seconds.

1 Answers

This can be done by creating a new column to group by that takes into account both the categorical and time information. First, make sure the dataframe is ordered by time:

df['time'] = pd.to_datetime(df['time'])
df = df.sort_values('time')

The wanted column can be created by using shift and cumsum:

df['group'] = (df['cat'].shift(1) != df['cat']).cumsum()

Intermediate result:

       cat                time  group
0       HP 2021-12-21 06:35:03      1
1       HP 2021-12-21 06:35:04      1
2  kellogs 2021-12-21 06:35:42      2
3  kellogs 2021-12-21 06:35:43      2
4       HP 2021-12-21 06:35:45      3

Now, we can use groupby and compute the number of seconds for each group:

df = df.groupby('group').agg( {'cat': 'first', 'time': ['first', 'last']})
df.columns = ["_".join(a) for a in df.columns.to_flat_index()]
df['time'] = df['time_last'] - df['time_first'] + pd.Timedelta(seconds=1)
df = df.rename(columns={'cat_first': 'cat'})

Finally, we sum up the number of seconds for each category:

df = df.groupby('cat')['time'].sum().reset_index()

Result:

       cat            time
0       HP 0 days 00:00:03
1  kellogs 0 days 00:00:02
Related