Pyspark Dataframe GroupBy Count is Different Every Time

Viewed 1472

I am new to pyspark so I feel like I'm missing something simple and stupid here but I am missing it. I read the contents of a parquet file into a spark dataframe. Then I groupby a binary column and count the results. The numbers are different every time. Is this an issue of syncing data between spark nodes or does this have something to do with the lazy execution or am I just missing some basic fundamental spark principle? I'm very confused by these results.

df = spark.read.parquet(input_file)
df = df.limit(2000)
print(df.count())
print(df.groupBy('STATUS').count().collect())
print(df.groupBy('STATUS').count().collect())
print(df.groupBy('STATUS').count().collect())

>>> 2000
>>> [Row(STATUS=0, count=1613), Row(STATUS=1, count=387)]
>>> [Row(STATUS=0, count=1528), Row(STATUS=1, count=472)]
>>> [Row(STATUS=0, count=1646), Row(STATUS=1, count=354)]

Below is the df schema:

root
 |-- GRP_ID: long (nullable = true)
 |-- WEK_ID: long (nullable = true)
 |-- WEK_BGN_DT: string (nullable = true)
 |-- WEK_END_DT: string (nullable = true)
 |-- FEATURES: vector (nullable = true)
 |-- STATUS: long (nullable = true)

I should also note that if I convert the spark dataframe to pandas and get a count, it works just fine:

dfp = df.toPandas()    
print(dfp['STATUS'][dfp['STATUS'] == 0].count())
print(dfp['STATUS'][dfp['STATUS'] == 1].count())
print(dfp['STATUS'][dfp['STATUS'] == 0].count())
print(dfp['STATUS'][dfp['STATUS'] == 1].count())
print(dfp['STATUS'][dfp['STATUS'] == 0].count())
print(dfp['STATUS'][dfp['STATUS'] == 1].count())

>>> 1494
>>> 506
>>> 1494
>>> 506
>>> 1494
>>> 506
0 Answers
Related