I'm working with a PySpark dataframe and I need to do a Binomial regression, with more than one trial in each row. For example, my table looks like this:
┌──────────┬──────────┬─────────────┬────────────┐ │ Features │ # Trials │ # Successes │ # Failures │ ├──────────┼──────────┼─────────────┼────────────┤ │ ... │ 10 │ 4 │ 6 │ │ ... │ 7 │ 2 │ 5 │ │ ... │ 5 │ 4 │ 1 │ └──────────┴──────────┴─────────────┴────────────┘
I don't want to 'ungroup' the data. In statsmodels, there is a possibility to directly do a Binomial Regression on the grouped data with a patsy formula:
formula = '# Successes + # Failures ~ Features'
Is there a way to do so in PySpark as well?