Sparse Gaussian Processes using Pseudo-inputs
Edward Snelson, Zoubin Ghahramani
Abstract
Edward Snelson, Zoubin Ghahramani
Abstract
We present a new Gaussian process (GP) regression model whose co-variance is parameterized by the the locations of M pseudo-input points, which we learn by a gradient based optimization. We take M ¿ N, where N is the number of real data points, and hence obtain a sparse regression method which has O(M2N) training cost and O(M2) pre-diction cost per test case. We also find hyperparameters of the covari-ance function in the same joint optimization. The method can be viewed as a Bayesian regression model with particular input dependent noise. The method turns out to be closely related to several other sparse GP ap-proaches, and we discuss the relation in detail. We finally demonstrate its performance on some large data sets, and make a direct comparison to other sparse GP methods. We show that our method can match full GP performance with small M, i.e. very sparse solutions, and it significantly outperforms other approaches in this regime. 1
OpenAlex reports 1314 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
We present a new Gaussian process (GP) regression model whose co-variance is parameterized by the the locations of M pseudo-input points, which we learn by a gradient based optimization. We take M ¿ N, where N is the number of real data points, and hence obtain a sparse regression method which has O(M2N) training cost and O(M2) pre-diction cost per test case. We also find hyperparameters of the covari-ance function in the same joint optimization. The method can be viewed as a Bayesian regression model with particular input dependent noise. The method turns out to be closely related to several other sparse GP ap-proaches, and we discuss the relation in detail. We finally demonstrate its performance on some large data sets, and make a direct comparison to other sparse GP methods. We show that our method can match full GP performance with small M, i.e. very sparse solutions, and it significantly outperforms other approaches in this regime. 1
Key concepts: Hyperparameter, Gaussian process, Bayesian optimization, Kriging, Parameterized complexity, Computer science, Covariance, Regression