Bayesian Linear Regression
Tom Minka
Abstract
Tom Minka
Abstract
This note derives the posterior, evidence, and predictive density for linear multivariate regression under zero-mean Gaussian noise. Many Bayesian texts, such as Box & Tiao (1973), cover linear regression. This note contributes to the discussion by paying careful attention to invariance issues, demonstrating model selection based on the evidence, and illustrating the shape of the predictive density. Piecewise regression and basis function regression are also discussed. 1 Introduction The data model is that an input vector x of length m multiplies a coefficient matrix A to produce an output vector y of length d, with Gaussian noise added: y = Ax + e (1) e N (0; V) (2) p(yjx; A;V) N (Ax; V) (3) This is a conditional model for y only: the distribution of x is not needed and in fact irrelevant to all inferences in this paper. As we shall see, conditional models create subtleties in Bayesian inference. In the special case x = 1 and m = 1, the conditioning disappears and we simply have a ...
OpenAlex reports 79 citations for this work. Citation counts describe recorded attention and do not establish research quality.
A contribution statement is not available in the OpenAlex record.
Method details are not available in the OpenAlex metadata.
Findings are not separately available in the OpenAlex metadata.
Limitations are not available in the OpenAlex metadata.
Application details are not available in the OpenAlex metadata.
This note derives the posterior, evidence, and predictive density for linear multivariate regression under zero-mean Gaussian noise. Many Bayesian texts, such as Box & Tiao (1973), cover linear regression. This note contributes to the discussion by paying careful attention to invariance issues, demonstrating model selection based on the evidence, and illustrating the shape of the predictive density. Piecewise regression and basis function regression are also discussed. 1 Introduction The data model is that an input vector x of length m multiplies a coefficient matrix A to produce an output vector y of length d, with Gaussian noise added: y = Ax + e (1) e N (0; V) (2) p(yjx; A;V) N (Ax; V) (3) This is a conditional model for y only: the distribution of x is not needed and in fact irrelevant to all inferences in this paper. As we shall see, conditional models create subtleties in Bayesian inference. In the special case x = 1 and m = 1, the conditioning disappears and we simply have a ...
Key concepts: Bayesian multivariate linear regression, Bayesian linear regression, Proper linear model, Segmented regression, Linear regression, Mathematics, Bayesian probability, Statistics