I have a subset of the dataframe
df = pd.DataFrame(
{
'id': ['1001','1002','1003','1004','1005','1006','1007','1008','1009','1010'],
'colA': ['H','L B','L H','L B','L S B','B','B S L','L B S','L S B','L S B'],
'colB': ['H','L|B','H|L','H|L','L|S|B','L|S|B','L|S|B','L|S|B','L|S','L']
}
)
I'm doing row level comparison for this dataframe. I want to check whether all the letters in row['colA'] match with all the letters in row['colB'], regardless of what order they appear and ignoring the | in colB. This is the logic for the function, but it doesn't work as intended and how do I update it to ignore |
def match_or_not(df):
for index,row in df.iterrows():
if row['colA'] == row['colB']:
print ("Match for "+str(row['id']))
else:
print ("Not match for "+str(row['id']))
I need help to update the condition following the if keyword in the above function, how can I write to get desire output. The cases for which it should match and not match are shown in the picture:
