01 / Start with uncertainty
Bayes updates a distribution
A measurement does not tell us exactly what is true. It tells us which possibilities have become more plausible.
Suppose we want to infer an unknown quantity . Before seeing the data, we describe what we believe with a prior, . A model describes how different values of could generate data : this is the likelihood, . Bayes’ rule combines them:
The posterior is our updated distribution over . The evidence, , normalises it so that its total probability is one. Evidence also matters when comparing models: it measures how well a model predicts the data after averaging over its prior uncertainty.
What was plausible before this measurement?
How compatible is this measurement with each possible parameter value?
What is plausible after combining the two?
A likelihood is not automatically a probability distribution over parameters. It describes the probability density of the data given parameters. Multiplying it by a prior, then normalising, gives the posterior.
Why can the evidence be difficult to calculate?
We need to average over all parameter values:
For some simple models this integral is available analytically. For a nonlinear model with many parameters, it can be expensive or intractable. This is one reason to approximate the posterior.
02 / Read the expression
What does mean?
It is a prediction-error score that takes uncertainty into account.
Let be the difference between the observed data and the model’s prediction. Write precision as (capital pi), rather than . Precision is inverse variance for a scalar, or inverse covariance for a vector:
Small variance means high precision. Under the model, a precise measurement is expected to be close to its prediction, so the same error counts more strongly. Low precision makes a given mismatch less surprising.
For a single error, the expression is just . If , a standard deviation of gives a score of . A standard deviation of gives a score of . The error has not changed; its meaning has changed because the expected uncertainty is different.
For a vector, is the same vector written as a row. Multiplication produces one scalar score. For independent errors, this reduces to a weighted sum:
Three expressions, three roles
- : the prediction error.
- : the precision-weighted error, still a vector.
- : the total weighted squared error, a scalar.
If errors are correlated, off-diagonal terms matter too. The score then asks whether the pattern of errors is surprising under the covariance model. Precision is an assumption about reliability; assigning high precision does not make a measurement trustworthy.
Where does the Gaussian distribution enter?
With , the negative log likelihood is
If the covariance is fixed, the last two terms do not affect the parameter gradient. If we estimate the noise covariance too, we must retain them: reducing precision cannot be treated as a cost-free way to make errors disappear.
03 / An exact Bayesian update
Move towards evidence, by the right amount
Suppose our prior is , and the measurement is , where . The posterior is also Gaussian. Its precision is the sum of the two precisions, and its mean is their weighted average:
The same result can be written as a correction to the old mean:
This is the connection: a Bayesian update uses a prediction error, but the amount we move depends on how precise the new evidence is relative to our existing belief.
How much should one measurement change your mind?
Move the sliders. The posterior is calculated analytically, with no fitting algorithm.
The observation moves the mean 80% of the way from 10 to 14. Posterior uncertainty is smaller than either starting uncertainty.
With the default values, the prior variance is and the measurement variance is . The gain is : the mean moves from to , and posterior variance is . Make the measurement noisier and the correction shrinks. This direct-observation Gaussian example is exact; more complicated models need more work.
The likelihood curve is shown as a density over the horizontal axis for comparison. In this particular direct-observation model it also integrates to one over θ; that does not hold for likelihoods in general.
04 / From a score to a learning rule
Differentiate the mismatch
The quadratic error is a cost. Its derivative tells us how to change a parameter to reduce that cost.
Start with the Gaussian data cost . Define the Jacobian : it tells us how each prediction changes when we change each parameter. Because , increasing a prediction decreases its error. The chain rule gives
Add a Gaussian prior with mean and precision . Ignoring constants that are fixed with respect to the parameters, the negative log posterior is
A gradient-descent step is therefore
The first term moves parameters to reduce the data mismatch. The second pulls them towards the prior. The step size controls the numerical optimisation: it is not, in general, the exact Bayesian gain from the previous example.
Why the factor of one half? Differentiating a square gives a factor of two. The half cancels it, leaving the clean precision-weighted error. The square and the update are related, but they are not the same expression.
05 / Fit a function
Would this reinvent variational Laplace?
It recovers part of the machinery. To reach variational Laplace, we also need an approximation to posterior uncertainty and a variational objective.
Imagine fitting a decaying response, . The amplitude and decay rate are unknown. Each observation creates an error, and the Jacobian maps those errors back to parameter changes.
Curvature can make the updates more efficient. A Gauss–Newton approximation to the curvature of the negative log posterior is
Instead of choosing one scalar step size, it rescales and couples the parameter corrections:
This uses the same information that helps describe local uncertainty: steep curvature implies a narrow range of plausible parameters; shallow curvature implies a wider range.
Fit an exponential response
Ten fixed synthetic observations. Take a step, or run the fit. Compare the prediction with the local uncertainty ellipse.
95% contour of the local Gaussian. A tilted ellipse means the parameter uncertainties are coupled.
Changing noise or prior precision restarts the fit. Priors: A = 1.2 ± 0.8, k = 0.25 ± 0.35 (mean ± SD, before the multiplier). The generating parameters are A = 2.5 and k = 0.65 s⁻¹. Steps use backtracking to reduce the objective.
Try this: run the fit, then increase the assumed noise and run it again. The observations have not changed, but they now exert less influence. Increase prior precision and the solution is pulled more strongly towards the prior. Switch to gradient descent to see how curvature affects the route to the solution.
This demonstration performs MAP optimisation and displays the inverse Gauss–Newton curvature. Before convergence, the ellipse is only a local curvature summary; near a valid optimum it approximates posterior uncertainty. It is not a full variational Laplace implementation.
06 / Fit a distribution
Variational inference asks a bigger question
Instead of asking only “which parameter value fits?”, ask “which tractable distribution best approximates the posterior?”
Choose a family of distributions , perhaps Gaussians described by a mean and covariance . We then optimise those distribution parameters. One common objective is the evidence lower bound (ELBO), also called variational free energy under the convention used here:
The first term is expected log likelihood: how well the data are explained across the uncertainty in . The second is the divergence from the prior: how much the distribution changes its beliefs to achieve that fit. It is often called complexity, but it is not simply the number of parameters.
For a nonlinear model, optimising the mean of need not give exactly the MAP estimate. The objective considers a spread of parameter values, rather than evaluating the fit at just one point.
The key identity is
KL divergence is non-negative. So is a lower bound on log evidence, and maximising brings closer to the posterior in this particular direction of KL divergence. Some texts define free energy with the opposite sign and minimise it; always check the convention.
Make a Gaussian approximate the posterior
This uses the prior and observation from Experiment 1. Adjust the mean and width of q. Its score rewards both the fit and the right uncertainty.
Here the exact posterior is known, so we can show the gap. For difficult models, that exact comparison is usually unavailable.
Try this: match the exact posterior, then make very narrow. Concentrating on the best-looking parameter value does not keep improving the bound. We lose the posterior’s uncertainty, and the divergence grows. A distribution’s spread is part of what inference must get right.
What is being calculated in this experiment?
For , the expected log likelihood includes both the squared error of its mean and its variance:
The prior divergence is
The exact evidence is . All scores use natural logarithms (nats). “Match exact posterior” sets the known analytical solution; it does not simulate an optimisation algorithm.
07 / Put the pieces together
From Laplace approximation to variational Laplace
Near a well-behaved posterior mode, approximate the negative log posterior by a quadratic:
Exponentiating a negative quadratic gives a Gaussian. This is the Laplace approximation:
Here is the positive-definite Hessian at the mode. In nonlinear least squares, the Gauss–Newton matrix is a useful approximation to that Hessian; it omits terms involving residuals and second derivatives of the predictions.
Variational Laplace uses Gaussian approximations and local curvature within a variational free-energy scheme. In implementations such as SPM’s model inversion, mean updates, posterior covariance, and noise-precision estimates work together. The details go beyond taking a MAP step and appending an error bar.
| Routine | What it gives you |
|---|---|
| Minimise squared prediction errors | A least-squares parameter estimate; with fixed Gaussian noise, a maximum-likelihood estimate. |
| Add a prior and optimise | A MAP estimate: one best parameter vector under the posterior density. |
| Approximate curvature at a valid mode | A local Gaussian posterior approximation: MAP plus approximate covariance. |
| Optimise a distribution using a variational bound | Variational inference; the approximation can use many different families. |
| Use a Gaussian/local-curvature scheme for variational inversion | Variational Laplace; typically with coordinated mean, covariance and precision updates. |
So, would a function-fitting routine be reinventing variational Laplace? It could be heading there. Precision-weighted errors, derivatives and curvature are shared ingredients. The defining extra step is treating inference as an approximation to a distribution, with uncertainty included in the variational objective.
Where the approximation can fail
A single Gaussian can miss multiple posterior modes, asymmetry or curved dependencies. A local optimum may not be the global one. The covariance is conditional on the model and noise assumptions, and an approximate evidence is not exact evidence. Check whether the approximation is adequate for the scientific question.
08 / Bring it back to a generative model
The same logic scales to neural models
In a neural generative model, might contain coupling strengths, time constants or other physiological parameters. A forward model predicts a signal from those parameters. Comparing the predicted and observed signals creates a vector of errors. The Jacobian connects changes in the signal to changes in the underlying parameters.
The logic remains: specify plausible parameters and a noise model; predict the observations; use precision to interpret the errors; update the parameter beliefs; and describe how much uncertainty remains. A good fit alone does not establish that a parameter is well identified or that the model is the right explanation.
What signal should these parameters generate?
Which parameter changes explain the mismatch?
Which combinations remain plausible after seeing the data?
A small notation guide
| Symbol | Meaning |
|---|---|
| , , | Observed data, model prediction, and their difference. |
| , | Covariance and its inverse, precision. Lowercase is scalar precision. |
| , | Prediction Jacobian and local curvature of the negative log posterior. |
| , | Prior mean and prior precision. |
| , , | Approximate posterior, its mean and its covariance. |
| , | Variational lower bound and Kullback–Leibler divergence. |
Follow the ideas further
CPNS student guidePlace inference within the wider modelling workflow. Interactive conceptsExplore related ideas in the CPNS teaching collection. CPNS methodsConnect the introduction to the lab’s methodological work.
Primary references
- Friston, K. J., Mattout, J., Trujillo-Barreto, N., Ashburner, J. & Penny, W. D. (2007). Variational free energy and the Laplace approximation. NeuroImage, 34, 220–234. Author-hosted PDF.
- Blei, D. M., Kucukelbir, A. & McAuliffe, J. D. (2017). Variational Inference: A Review for Statisticians. Journal of the American Statistical Association, 112, 859–877.
- SPM: spm_nlsi_GN.m. A concrete implementation to read after understanding the distinction between MAP fitting and variational inversion.
The examples are deliberately small. The first and third have exact Gaussian answers; the nonlinear fit uses an approximation. Keeping those cases separate makes the connection between Bayes, optimisation and variational inference easier to see.