I am performing linear regression using Sklearn and statsmodels.
I know that Sklearn and statsmodels produce the same result. As shown below, Sklearn and statsmodels made the same result, but the results were different even though the coefficients were the same when the intercept is zero using fit_intercept=False in Sklearn.
Can you explain the reason? Or give me any method when I use fit_intercept=False in Sklearn.
import numpy as np
import statsmodels.api as sm
from sklearn.linear_model import LinearRegression
# dummy data:
y = np.array([1,3,4,5,2,3,4])
X = np.array(range(1,8)).reshape(-1,1) # reshape to column
# intercept is not zero : the result are the same
# scikit-learn:
lr = LinearRegression()
lr.fit(X,y)
print(lr.score(X,y))
# 0.16118421052631582
# statsmodels
X_ = sm.add_constant(X)
model = sm.OLS(y,X_)
results = model.fit()
print(results.rsquared)
# 0.16118421052631582
# intercept is zero : the result are different
# scikit-learn:
lr = LinearRegression(fit_intercept=False)
lr.fit(X,y)
print(lr.score(X,y))
# -0.4309210526315792
# statsmodels
model = sm.OLS(y,X)
results = model.fit()
print(results.rsquared)
# 0.8058035714285714