Python normalised linear model produces odd results

Viewed 40

Whenever I fit a linear model where the features have been normalised (0 mean, 1 std), if I then use the model to score/predict new data where all of the features are all set to 0, I get a non-zero answer and do not understand why.

I have created a toy example below to illustrate what I mean.

Import libraries

import matplotlib.pyplot as plt
import numpy as np
import pandas as pd
from scipy.optimize import lsq_linear
from sklearn.preprocessing import StandardScaler

Generate some example data

x = pd.DataFrame({
    "a": [1,2,4,5,6,7],
    "b": [2,5,3,5,6,9]
})

noise = abs(np.random.randn(x.shape[0]))
y = x['a'] + x['b'] + noise

Normalise the data

# normalise x
x_scaler = StandardScaler()
x_norm = pd.DataFrame(x_scaler.fit_transform(x))

# normalise y
y_mean = np.mean(y)
y_std = np.std(y)
y_norm = (y - y_mean) / y_std

Fit linear model

model = lsq_linear(A=x_norm, b=y_norm, verbose=0)

# reconstruct fit (for checking
coefficients = model['x']
fit_norm = x_norm.dot(coefficients).reset_index(drop=True)
fit = (fit_norm * y_std) + y_mean

# check fits
plt.plot(y)
plt.plot(fit)
plt.title("fit - original data")

plt.plot(y_norm)
plt.plot(fit_norm)
plt.title("fit - normalised data")

Both fits look good/reasonable

Now, for the part which I do not understand

Score model using features all set to 0

test = x.copy()

# Set features to 0
test[test.columns] = 0

# Normalise the data (using the same mean and std from before)
test_norm = pd.DataFrame(x_scaler.transform(test))

# Generate a normalised fit
test_fit_norm = test_norm.dot(coefficients).reset_index(drop=True)

# Un-normalise the fit
test_fit = (test_fit_norm * y_std) + y_mean
print(test_fit)

Does anyone know why the value of test_fit is non-zero? I've noticed it only happens if the linear model has residuals/a non-perfect fit, but I still do not understand why.

0 Answers
Related