I have a dataframe:
df_test = pd.DataFrame({'col': ['paris', 'paris', 'nantes', 'berlin', 'berlin', 'berlin', 'tokyo'],
'id_res': [12, 12, 14, 28, 8, 4, 89]})
col id_res
0 paris 12
1 paris 12
2 nantes 14
3 berlin 28
4 berlin 8
5 berlin 4
6 tokyo 89
I want to create a "check" column whose values are as follows:
- If a value in "col" has a duplicate and these duplicates have the same id_res, the value of "check" is False for duplicates
- If a value in "col" has duplicates and the "id_res" of these duplicates are different, assign True in "check" for the largest "id_res" value and False for the smallest
- If a value in "col" has no duplicates, the value of "check" is False.
The output I want is therefore:
col id_res check
0 paris 12 False
1 paris 12 False
2 nantes 14 False
3 berlin 28 True
4 berlin 8 False
5 berlin 4 False
6 tokyo 89 False
I tried with groupby but no satisfactory result. Can anyone help me plz