from sklearn.datasets import make_blobs
from sklearn.decomposition import PCA
SEED = 123
X, y = make_blobs(n_samples=1000, n_features=5000, cluster_std=90., random_state=SEED)
pca = PCA(2)
pca.fit(X)
pca1, pca2 = pca.components_
pcaX = pca.transform(X)
pcaXnp = np.array([X @ pca1, X @ pca2]).T
And if you print out pcaX and pcaXnp you'll see that they're similar but that they don't agree with each other. Why should these differ? It seems like ".components_" should return what sklearn is going to multiply the matrix by, is there a reason why it's just an approximation of what the multiplication will be?