Source of small deviance between SAS PROC LOGISTIC and Python sklearn LogisticRegression (unpenalized)?

Viewed 612

I'm trying to match the results of SAS's PROC LOGISTIC with sklearn in Python 3. SAS uses unpenalized regression, which I can achieve in sklearn.linear_model.LogisticRegression with the option C = 1e9 or penalty='none'.

That should be the end of the story, but I still notice a small difference when I use a public data set from UCLA and try to replicate their multiple regression of FEMALE and MATH on hiwrite.

This is my Python script:

# module imports
import pandas as pd
from sklearn.linear_model import LogisticRegression

# read in the data
df = pd.read_sas("~/Downloads/hsb2-4.sas7bdat")

# FE
df["hiwrite"] = df["write"] >= 52

print("\n\n")
print("Multiple Regression of Female and Math on hiwrite:")
feature_cols = ['female','math']

y=df["hiwrite"]
X=df[feature_cols]

# sklearn output
model = LogisticRegression(fit_intercept = True, C = 1e9)
mdl = model.fit(X, y)
print(mdl.intercept_)
print(mdl.coef_)

which yields:

Multiple Regression of Female and Math on hiwrite:
[-10.36619688]
[[1.63062846 0.1978864 ]]

UCLA has this result from SAS:

             Analysis of Maximum Likelihood Estimates
                               Standard          Wald
Parameter    DF    Estimate       Error    Chi-Square    Pr > ChiSq
Intercept     1    -10.3651      1.5535       44.5153        <.0001
FEMALE        1      1.6304      0.4052       16.1922        <.0001
MATH          1      0.1979      0.0293       45.5559        <.0001

which is close, but as you can see the intercept parameter estimate is different at the 3rd decimal place and the estimate on female is different at the 4th decimal place. I tried changing some of the other parameters (like tol and max_iter as well as the solver) but it did not change the results. I also tried the Logit in statsmodel.api - it matches sklearn, not SAS. R matches Python on the intercept and first coefficient, but is slightly different from both SAS and Python on the second coefficient...

Update: I went to SAS' community to look for answers and someone mentioned that it may be due to differences in the iterative maximum likelihood algorithm convergence. I feel like that sounds right though I've already tried to get at that with the available options in Python.

Any thoughts on the source of the error and how to make Python match SAS?

0 Answers
Related