I have a spark dataframe like below. (just an example. My real data has millions of rows):
df = pd.DataFrame({'ZIP1': ['50069', '50069', '50704', '50704', '52403', '52403'],
'ZIP2': ['50704', '52403', '50069', '52403', '50069', '50704'],
'STATE': ['IA', 'IA', 'IA', 'IA', 'IA', 'IA'],
'REGION': ['MIDWEST', 'MIDWEST', 'MIDWEST', 'MIDWEST', 'MIDWEST', 'MIDWEST'] } )
sdf = spark.createDataFrame(df)
ZIP1 ZIP2 STATE REGION
0 50069 50704 IA MIDWEST
1 50069 52403 IA MIDWEST
2 50704 50069 IA MIDWEST
3 50704 52403 IA MIDWEST
4 52403 50069 IA MIDWEST
5 52403 50704 IA MIDWEST
If two zipcodes from ZIP1 and ZIP2 columns are the same combination, I need to remove one row. For example, row 0 and row 2, the zipcodes are simply the same combination, but in reversed order. I need to remove either row 0 or row 2. Likewise, remove either row 1 or row 4....
Does anyone know how to achieve this in pyspark? Pyspark solution is needed. If someone can provide solutions in both pyspark and python, that is a plus. Thanks!