I have a dataframe with ID and some email addresses
personid sup1_email sup2_email sup3_email sup4_email
1 evan.o@abc.com jon.k@abc.com kelm.q@abc.com john.d@abc.com
5 evan.o@abc.com polly.u@abc.com jim.e@ABC.COM nan
11 jim.y@abc.com manfred.a@abc.com greg.s@Abc.com adele.a@abc.com
52 jim.y@abc.com manfred.a@abc.com greg.s@Abc.com adele.a@abc.com
65 evan.o@abc.com lenny.t@yahoo.com john.s@abc.com sally.j@ABC.com
89 dom.q@ABC.com laurie.g@Abc.com topher.u@abc.com ross.k@qqpower.com
I would like to locate the rows which do not match the list of accepted email values (ie NOT '@abc.com', '@ABC.COM', '@Abc.com'). What i'd like to get is this
personid sup1_email sup2_email sup3_email sup4_email
65 evan.o@abc.com lenny.t@yahoo.com john.s@abc.com sally.j@ABC.com
89 dom.q@ABC.com laurie.g@Abc.com topher.u@abc.com ross.k@qqpower.com
I've written the following code and it works but I have to manually check for each sup_email column and repeat the process, which is inefficient
#list down all the variations of accepted email domains
email_adds = ['@abc.com','@ABC.COM','@Abc.com']
#combine the variations of email addresses in the list
accepted_emails = '|'.join(email_adds)
not_accepted = df.loc[~df['sup1_email'].str.contains(accepted_emails, na=False)]
I was wondering if there was a more efficient way to do this using a for loop. What i've managed so far is to show a column which contains a non-accepted email, but it is not showing the rows that contain non-accepted emails. Appreciate any form of help I can get, thank you.
sup_emails = df[['sup1_email','sup2_email', 'sup3_email', 'sup4_email']]
#for each sup column, check if the accepted email addresses are not in it
for col in sup_emails:
if any(x not in col for x in accepted_emails):
print(col)