I have the below dataframe:
x = pd.DataFrame({
"item" : ["a", "a", "a", "b", "c", "c"],
"vote" : [1, 0, 1, 1, 0, 0],
"timestamp" : ["2020-06-07 11:04:26", "2020-06-07 11:03:37", "2020-06-07 11:09:18", "2020-06-07 11:04:40", "2020-06-07 11:09:11", "2020-06-07 11:09:23"]
})
item vote timestamp
a 1 2020-06-07 11:04:26
a 0 2020-06-07 11:03:37
a 1 2020-06-07 11:09:18
b 1 2020-06-07 11:04:40
c 0 2020-06-07 11:09:11
c 0 2020-06-07 11:09:23
How do I drop_duplicates on item column, and use the timestamp column as a tiebreaker: keep the latest one?
The final dataframe should look like this:
item vote timestamp
a 1 2020-06-07 11:09:18
b 1 2020-06-07 11:04:40
c 0 2020-06-07 11:09:23
