Remove duplicates in a row pandas

Viewed 41

I have a df

Name  Symbol              Dummy
A     (BO),(BO),(AD),(TR)   2
B     (TV),(TV),(TV)        2
C     (HY)                  2
D     (UI)                  2

I need df as

Name  Symbol              Dummy
A     (BO),(AD),(TR)        2
B     (TV)                  2
C     (HY)                  2
D     (UI)                  2

Tried with this function but not working as expected.

drop_duplicates
2 Answers

Split the strings around delimiter ,, then dedupe using dict.fromkeys which also preserves the order of strings, finally join around delimiter ,

df['Symbol'] = df['Symbol'].str.split(',').map(dict.fromkeys).str.join(',')

  Name          Symbol  Dummy
0    A  (BO),(AD),(TR)      2
1    B            (TV)      2
2    C            (HY)      2
3    D            (UI)      2

Another method

#original DF

index col1 col2
0 (BO),(BO),(AD),(TR) 2
df.col1 = df.col1.str.split(',').apply(lambda x: sorted(set(x), key=x.index)).str.join(',')
df

#output

index col1 col2
0 (BO),(AD),(TR) 2

If values order not important you can simply do:

df.col1 = df.col1.str.split(',').apply(lambda x: set(x)).str.join(',')
df

#output

index col1 col2
0 (AD),(BO),(TR) 2
Related