I'd like to apply a function to two columns (A and B) of a pandas dataframe that tests if each of their values match the same result in a dictionary. I'd like it to return the result to a third column.
I've tried the code below and close variants but I keep getting errors and I think there's something fundemental that I'm not understanding about the data structure. Can anyone explain where am I going wrong? I can imagine cumbersome alternative ways to do this but I'm sure there must be an elegant solution.
def do_they_match(A1,A2):
if A1 in dictionary and A2 in dictionary and dictionary[A1] == dictionary[A2]:
return 1
else:
return 0
df['match'] = df.apply(lambda x: do_they_match(x['A'],x['B']))
## also tried ##
df = df.assign(link=lambda x: do_they_match(x['A'],x['B']))
For context, the errors I get are IndexError: ('A', 'occurred at index A') or TypeError: 'Series' objects are mutable, thus they cannot be hashed for the alternative code on the last line.The values in both dataframe columns and in the dictionary are all strings.
Thanks for the help!