Whenever I fit a linear model where the features have been normalised (0 mean, 1 std), if I then use the model to score/predict new data where all of the features are all set to 0, I get a non-zero answer and do not understand why.
I have created a toy example below to illustrate what I mean.
Import libraries
import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
from scipy.optimize import lsq_linear
from sklearn.preprocessing import StandardScaler
Generate some example data
x = pd.DataFrame({
"a": [1,2,4,5,6,7],
"b": [2,5,3,5,6,9]
})
noise = abs(np.random.randn(x.shape[0]))
y = x['a'] + x['b'] + noise
Normalise the data
# normalise x
x_scaler = StandardScaler()
x_norm = pd.DataFrame(x_scaler.fit_transform(x))
# normalise y
y_mean = np.mean(y)
y_std = np.std(y)
y_norm = (y - y_mean) / y_std
Fit linear model
model = lsq_linear(A=x_norm, b=y_norm, verbose=0)
# reconstruct fit (for checking
coefficients = model['x']
fit_norm = x_norm.dot(coefficients).reset_index(drop=True)
fit = (fit_norm * y_std) + y_mean
# check fits
plt.plot(y)
plt.plot(fit)
plt.title("fit - original data")
plt.plot(y_norm)
plt.plot(fit_norm)
plt.title("fit - normalised data")
Both fits look good/reasonable
Now, for the part which I do not understand
Score model using features all set to 0
test = x.copy()
# Set features to 0
test[test.columns] = 0
# Normalise the data (using the same mean and std from before)
test_norm = pd.DataFrame(x_scaler.transform(test))
# Generate a normalised fit
test_fit_norm = test_norm.dot(coefficients).reset_index(drop=True)
# Un-normalise the fit
test_fit = (test_fit_norm * y_std) + y_mean
print(test_fit)
Does anyone know why the value of test_fit is non-zero? I've noticed it only happens if the linear model has residuals/a non-perfect fit, but I still do not understand why.