Will this work for you? The key is to sort the data, then apply df.duplicated(), which has very high efficiency rather than looping through each record like .apply(lambda) functions
import pandas as pd
import numpy as np
df = pd.DataFrame({
'ID': [39203920, 32323, 22222, 392999],
'EmailAddress': ['john@gmail.com', 'j@email.com', 'j@email.com', 'john@gmail.com'],
'Name': ['John', np.nan, 'Jane', 'John'],
'Country': ['UK', 'UK', 'UK', 'UK'],
'Distance': [12, 12, 12, 12],
'IDLen': [8, 5, 5, 6],
'NonNAN': [6, 5, 6, 6] })
df = df.sort_values(['EmailAddress', 'NonNAN', 'IDLen'], ascending=[True, False, True])
ID EmailAddress Name Country Distance IDLen NonNAN
2 22222 j@email.com Jane UK 12 5 6
1 32323 j@email.com NaN UK 12 5 5
3 392999 john@gmail.com John UK 12 6 6
0 39203920 john@gmail.com John UK 12 8 6
Based on your rules, I have sorted the data so that the desired record is located first. When df.duplicated() is applied on EmailAddress, the first record will be kept
df1 = df[~df.duplicated('EmailAddress')]
ID EmailAddress Name Country Distance IDLen NonNAN
2 22222 j@email.com Jane UK 12 5 6
3 392999 john@gmail.com John UK 12 6 6
df2 = df[df.duplicated('EmailAddress')]
ID EmailAddress Name Country Distance IDLen NonNAN
1 32323 j@email.com NaN UK 12 5 5
0 39203920 john@gmail.com John UK 12 8 6
If your ID column is numerical (ie, not alphanumeric), you can sort based on ascending ID, and there is no need for the column IDLen (because you would like the shortest one if 'NonNAN' is the same)