Why is my regressor predicting counts much further than the actual counts?

Viewed 174

I'm building a regression model (statsmodels.discrete.count_model.ZeroInflatedPoisson) in which my goal is to predict the "count" variable, as you can see below. Any suggestion on how I can improve this model?

#regression expression in Patsy notation. count depends on all these columns
expr = 'count ~ day_of_week + business_day + duration + Distance_KM + wind_speed + wind_deg + wind_gust + rain_1h + rain_3h + clouds_all + End_Station_Region_cat'

y_train, X_train = dmatrices(expr, df, return_type='dataframe')
y_test, X_test = dmatrices(expr, df, return_type='dataframe')

import statsmodels.discrete.count_model as cm

zip_training_results = cm.ZeroInflatedPoisson(endog=y_train, exog=X_train, exog_infl=X_train, inflation='logit', maxiter=1000, maxfun=500).fit_regularized()

Optimization terminated successfully    (Exit mode 0)
            Current function value: 0.030743127487834265
            Iterations: 184
            Function evaluations: 230
            Gradient evaluations: 184
    import statsmodels.discrete.count_model as cm



print(zip_training_results.summary())
                     ZeroInflatedPoisson Regression Results                    
===============================================================================
Dep. Variable:                   count   No. Observations:              3099093
Model:             ZeroInflatedPoisson   Df Residuals:                  3099081
Method:                            MLE   Df Model:                           11
Date:                 Sun, 04 Sep 2022   Pseudo R-squ.:                  0.6252
Time:                         17:17:42   Log-Likelihood:                -95276.
converged:                        True   LL-Null:                   -2.5418e+05
Covariance Type:             nonrobust   LLR p-value:                     0.000
==================================================================================================
                                     coef    std err          z      P>|z|      [0.025      0.975]
--------------------------------------------------------------------------------------------------
inflate_Intercept                 29.2202    488.934      0.060      0.952    -929.073     987.514
inflate_day_of_week                7.5087     11.894      0.631      0.528     -15.802      30.820
inflate_business_day              20.5779     59.774      0.344      0.731     -96.577     137.733
inflate_duration                 -85.6092    487.673     -0.176      0.861   -1041.431     870.212
inflate_Distance_KM               -1.9937     46.122     -0.043      0.966     -92.391      88.404
inflate_wind_speed                -1.0937      9.175     -0.119      0.905     -19.077      16.889
inflate_wind_deg                  -5.6348     48.297     -0.117      0.907    -100.296      89.026
inflate_wind_gust                  0.8242    763.772      0.001      0.999   -1496.142    1497.790
inflate_rain_1h                    0.5288     59.935      0.009      0.993    -116.942     118.000
inflate_rain_3h                    0.9703   7.97e+04   1.22e-05      1.000   -1.56e+05    1.56e+05
inflate_clouds_all                -0.1600      5.526     -0.029      0.977     -10.991      10.671
inflate_End_Station_Region_cat    -0.3111      2.362     -0.132      0.895      -4.940       4.318
Intercept                         -0.5539      0.020    -28.229      0.000      -0.592      -0.515
day_of_week                        0.0115      0.003      3.563      0.000       0.005       0.018
business_day                       0.0560      0.014      4.127      0.000       0.029       0.083
duration                           0.0149      0.000     80.515      0.000       0.015       0.015
Distance_KM                        0.0972      0.002     49.338      0.000       0.093       0.101
wind_speed                        -0.9573      0.009   -109.888      0.000      -0.974      -0.940
wind_deg                          -0.0106   8.82e-05   -119.945      0.000      -0.011      -0.010
wind_gust                          0.0001      0.015      0.008      0.994      -0.029       0.029
rain_1h                           -0.1926      0.019    -10.084      0.000      -0.230      -0.155
rain_3h                           -0.0724      0.026     -2.755      0.006      -0.124      -0.021
clouds_all                         0.0006      0.000      5.288      0.000       0.000       0.001
End_Station_Region_cat             0.0034      0.005      0.693      0.489      -0.006       0.013
==================================================================================================


zip_predictions = zip_training_results.predict(X_test,exog_infl=X_test)
predicted_counts=np.round(zip_predictions)
actual_counts = y_test['count']
print('ZIP RMSE='+str(np.sqrt(np.sum(np.power(np.subtract(predicted_counts,actual_counts),2)))))

ZIP RMSE=195.05127530985283

With the following image, I believe it is clear how the model presented in the code above ends up making predictions of counts much higher than the observed values. When Y is a low value, the regressor can make good predictions, however, what I would like to fix is that it predicts much higher values ​​for Y than the actual values.

plt.clf()
plt.hist([actual_counts, predicted_counts], log=True)
plt.legend(('orig','pred'))
plt.show()

enter image description here

Just Poisson attempt:

pipeline = Pipeline([('model', PoissonRegressor())])
pipeline.fit(X_train, y_train)
y_pred = pipeline.predict(X_test)
r2_test = metrics.r2_score(y_test, y_pred)

r2_test
-0.18668761012669255

    
    y_pred_train = pipeline.predict(X_train)
    r2_train = metrics.r2_score(y_train, y_pred_train)
    r2_train
-0.10978290552023906
0 Answers
Related