Posterior samples from fitted Gaussian process don't resemble predicted mean

Viewed 164

I am fitting a Gaussian process based on the second plot shown here and draw samples from it. However, the drawn samples do not resemble the fitted function and generally looks very different from the predicted mean (specifically, they are not as smooth and have sharp changes) and don't go through (or close) to the given data points (red dots in the plot).

Example plot (black line is predicted mean, blue and orange lines are the samples): enter image description here

After many runs, the result is always similar (even if not exactly the same). Any ideas what causes that and how I can make the drawn samples more similar to the mean?

The code used for generating the plot is

import numpy as np

from matplotlib import pyplot as plt

from sklearn.gaussian_process import GaussianProcessRegressor
from sklearn.gaussian_process.kernels import RBF, WhiteKernel as White


rng = np.random.RandomState(0)
X = rng.uniform(0, 5, 20)[:, np.newaxis]
y = 0.5 * np.sin(3 * X[:, 0]) + rng.normal(0, 0.5, X.shape[0])

kernel = 1.0 * RBF(length_scale=1.0, length_scale_bounds=(1e-2, 1e3)) \
    + White(noise_level=1e-5, noise_level_bounds=(1e-10, 1e+1))
gp = GaussianProcessRegressor(kernel=kernel, alpha=0.0)
gp.fit(X, y)

X_ = np.linspace(0, 5, 100)
y_mean, y_cov = gp.predict(X_[:, np.newaxis], return_cov=True)

plt.figure(figsize=(14,7))
plt.plot(X_, y_mean, 'k', lw=3, zorder=9)
plt.fill_between(X_, y_mean - np.sqrt(np.diag(y_cov)),
                 y_mean + np.sqrt(np.diag(y_cov)),
                 alpha=0.5, color='k')
plt.scatter(X[:, 0], y, c='r', s=50, zorder=10, edgecolors=(0, 0, 0))
plt.title("Initial: %s\nOptimum: %s\nLog-Marginal-Likelihood: %s"
          % (kernel, gp.kernel_, gp.log_marginal_likelihood(gp.kernel_.theta)))
y_samples = gp.sample_y(X_.reshape(-1, 1), 2)
plt.plot(X_, y_samples, lw=2)
plt.tight_layout()
1 Answers

I have been learning GP process recently and try to set an intuitive explanation:

Background:

The true (toy) function to learn is a sine function: (0.5 * sine x) + Normal Distribution adjustment (mean=0, SD = 0.5)

The Gaussian Process Regressor tries to learn and fit (based on some given learning pairs (X, y) / prior distribution) and successively arrive at some posterior distribution. As more learning points are given, usually the better is the prediction quality.

Current Code and Demo

The GP regressor will try to give a Predicted Mean + Confidence Range (eg. one SD in this case). In the case of this given sine related function, its true shape is typically like the Red line given my third attachment (note the red line constructed by 100 test points at times may deviate a lot from the others partly due to the random normal adjustment to the Sine equation.

I attach three more cases i) 80 prior learning points and ii) 10 prior learning points, iii) with a plot of true posterior function for 100 X inputs to compare with the current case of 20 points (pls note the 10 learning points plot show overfit in certain region, and underfit in others)

80 learning points

10 learning points

[True Function vs Learned3 What we should expect from a sample given by gp.sample_y ?

Our prediction should mainly be guided by the Predicted Mean (black) and the confidence level (grey). What you sample from the sample_y is a "One of the Many" Realization of the process. Intuitively it is like you are throwing a dice of N faces over a T time period, and as you can further subdivide T and subdivide N, this become infinitive combination. Even we keep a range and keep everything integers, it is still a huge combination of possible sub-distributions (samples).

So whether in prior or posterior distribution, i) those samples are for sure NOT as smooth as the aggregate (just like any individual games) and ii) there are many of them each behave all differently from each other but iii) aggregately they form certain Gaussian pattern.

Within my limited understanding, the sample is just a possible occurrence out of many, and most of the time, what turn out is within or close to the range of confidence, and the prediction generally is guided by the Predicted Mean (black line)

Reference: https://peterroelants.github.io/posts/gaussian-process-tutorial/

Related