I'm building a regression model (statsmodels.discrete.count_model.ZeroInflatedPoisson) in which my goal is to predict the "count" variable, as you can see below. Any suggestion on how I can improve this model?
#regression expression in Patsy notation. count depends on all these columns
expr = 'count ~ day_of_week + business_day + duration + Distance_KM + wind_speed + wind_deg + wind_gust + rain_1h + rain_3h + clouds_all + End_Station_Region_cat'
y_train, X_train = dmatrices(expr, df, return_type='dataframe')
y_test, X_test = dmatrices(expr, df, return_type='dataframe')
import statsmodels.discrete.count_model as cm
zip_training_results = cm.ZeroInflatedPoisson(endog=y_train, exog=X_train, exog_infl=X_train, inflation='logit', maxiter=1000, maxfun=500).fit_regularized()
Optimization terminated successfully (Exit mode 0)
Current function value: 0.030743127487834265
Iterations: 184
Function evaluations: 230
Gradient evaluations: 184
import statsmodels.discrete.count_model as cm
print(zip_training_results.summary())
ZeroInflatedPoisson Regression Results
===============================================================================
Dep. Variable: count No. Observations: 3099093
Model: ZeroInflatedPoisson Df Residuals: 3099081
Method: MLE Df Model: 11
Date: Sun, 04 Sep 2022 Pseudo R-squ.: 0.6252
Time: 17:17:42 Log-Likelihood: -95276.
converged: True LL-Null: -2.5418e+05
Covariance Type: nonrobust LLR p-value: 0.000
==================================================================================================
coef std err z P>|z| [0.025 0.975]
--------------------------------------------------------------------------------------------------
inflate_Intercept 29.2202 488.934 0.060 0.952 -929.073 987.514
inflate_day_of_week 7.5087 11.894 0.631 0.528 -15.802 30.820
inflate_business_day 20.5779 59.774 0.344 0.731 -96.577 137.733
inflate_duration -85.6092 487.673 -0.176 0.861 -1041.431 870.212
inflate_Distance_KM -1.9937 46.122 -0.043 0.966 -92.391 88.404
inflate_wind_speed -1.0937 9.175 -0.119 0.905 -19.077 16.889
inflate_wind_deg -5.6348 48.297 -0.117 0.907 -100.296 89.026
inflate_wind_gust 0.8242 763.772 0.001 0.999 -1496.142 1497.790
inflate_rain_1h 0.5288 59.935 0.009 0.993 -116.942 118.000
inflate_rain_3h 0.9703 7.97e+04 1.22e-05 1.000 -1.56e+05 1.56e+05
inflate_clouds_all -0.1600 5.526 -0.029 0.977 -10.991 10.671
inflate_End_Station_Region_cat -0.3111 2.362 -0.132 0.895 -4.940 4.318
Intercept -0.5539 0.020 -28.229 0.000 -0.592 -0.515
day_of_week 0.0115 0.003 3.563 0.000 0.005 0.018
business_day 0.0560 0.014 4.127 0.000 0.029 0.083
duration 0.0149 0.000 80.515 0.000 0.015 0.015
Distance_KM 0.0972 0.002 49.338 0.000 0.093 0.101
wind_speed -0.9573 0.009 -109.888 0.000 -0.974 -0.940
wind_deg -0.0106 8.82e-05 -119.945 0.000 -0.011 -0.010
wind_gust 0.0001 0.015 0.008 0.994 -0.029 0.029
rain_1h -0.1926 0.019 -10.084 0.000 -0.230 -0.155
rain_3h -0.0724 0.026 -2.755 0.006 -0.124 -0.021
clouds_all 0.0006 0.000 5.288 0.000 0.000 0.001
End_Station_Region_cat 0.0034 0.005 0.693 0.489 -0.006 0.013
==================================================================================================
zip_predictions = zip_training_results.predict(X_test,exog_infl=X_test)
predicted_counts=np.round(zip_predictions)
actual_counts = y_test['count']
print('ZIP RMSE='+str(np.sqrt(np.sum(np.power(np.subtract(predicted_counts,actual_counts),2)))))
ZIP RMSE=195.05127530985283
With the following image, I believe it is clear how the model presented in the code above ends up making predictions of counts much higher than the observed values. When Y is a low value, the regressor can make good predictions, however, what I would like to fix is that it predicts much higher values for Y than the actual values.
plt.clf()
plt.hist([actual_counts, predicted_counts], log=True)
plt.legend(('orig','pred'))
plt.show()
Just Poisson attempt:
pipeline = Pipeline([('model', PoissonRegressor())])
pipeline.fit(X_train, y_train)
y_pred = pipeline.predict(X_test)
r2_test = metrics.r2_score(y_test, y_pred)
r2_test
-0.18668761012669255
y_pred_train = pipeline.predict(X_train)
r2_train = metrics.r2_score(y_train, y_pred_train)
r2_train
-0.10978290552023906
