I'm using the PCA implementation from sklearn and wanted to export the loadings from the fitted model so I could transform anywhere else without using python.
However, I came up with an issue when I tried to validate that the dot product of the dataset with the loadings didn't give the same result as the transform function. Here's an example:
df = pd.DataFrame({'col1': [5,3,1,1,2,2,3,3,3],
'col2': [5,3,1,2,2,3,4,5,5],
'col3': [3,3,1,1,1,1,1,1,1]})
I import PCA from sklearn and fit the model with one component
from sklearn.decomposition import PCA
model = PCA(n_components=1)
model.fit(df)
With the fitted model, I transform the 3 columns dataset into one column
print(model.transform(df))
array([[ 3.13985669],
[ 0.4068059 ],
[-2.81207381],
[-2.0684094 ],
[-1.44554842],
[-0.701884 ],
[ 0.6646414 ],
[ 1.40830581],
[ 1.40830581]])
According to the sklearn docs I can access de loading in the components_ attribute. When I transform the dataset using the loadings I get a different output.
print(df.dot(model.components_.T).values)
array([[7.56137036],
[4.82831957],
[1.60943986],
[2.35310427],
[2.97596525],
[3.71962967],
[5.08615507],
[5.82981948],
[5.82981948]])
However, the difference between bot output seems to be constant
print(model.transform(df) - df.dot(model.components_.T).values)
[[-4.42151367]
[-4.42151367]
[-4.42151367]
[-4.42151367]
[-4.42151367]
[-4.42151367]
[-4.42151367]
[-4.42151367]
[-4.42151367]]
I was taught that PCA doesn't have intercept, but does this mean that the PCA implementation in sklearn includes the intercept? If so, is there a way to access this intercept without calling the difference between the transform funcion and the dot product of the data with the loadings?
Note: I know data normalization solves the problem of the intercept but I can't use it in this situation.