I'm trying to clean a list, by removing duplicates. For example:
bb = ['Gppe (Aspirin Combined)',
'Gppe Cap (Migraine)',
'Gppe Tab',
'Abilify',
'Abilify Maintena',
'Abstem',
'Abstral']
Ideally, I need to get the following list:
bb = ['Gppe',
'Abilify',
'Abstem',
'Abstral']
What I tried:
Split the list and remove duplicates (a naive approach)
list(set(sorted([j for bb_i in bb for j in bb_i.split(' ')])))
which leaves a lot of 'rubbish':
['(Aspirin',
'(Migraine)',
'Abilify',
'Abstem',
'Abstral',
'Cap',
'Combined)',
'Gppe',
'Maintena',
'Tab']
- Find the most frequent word:
Counter(['Gppe (Aspirin Combined)', 'Gppe Cap (Migraine)', 'Gppe Tab').most_common(1)[0][0]
But I'm not sure how to find similar words (a group)??
I am wondering, whether one can use a kind of 'groupby()' and first group by names and then remove duplicates within those names.