How to perform Tukey HSD test in Spark Dataframe?

Viewed 85

I am trying to perform the Tukey's test on a very large dataset using pyspark. Now I know in python we can use the pairwise_tukeyhsd library from the statsmodels.stats.multicomp module. That would require me to convert my spark data frame to pandas data frame which defeats the purpose of using RDD and will not work on my large dataset.

The other way is to manually do the test mathematically on the spark dataframes which is simple enough as shown here However, to compare the means with the Q_crit value, I would need Tukey's critical value table.

Is there any way to calculate the critical values on the Tukey table?

0 Answers
Related