MCMC (Markov Chain Monte Carlo)#
Sampling from a posterior via a Markov chain whose stationary distribution is that posterior.
Important
✨ AI-generated content. This page was written with the assistance of an AI language model and is provided as a learning aid. Despite careful review, it may still contain mistakes, omissions, or out-of-date information. Whether you are new to the topic, a team lead, or a senior practitioner, treat it as a starting point rather than an authoritative reference: read it critically and independently verify anything you act on (code, commands, figures, and factual claims) against official documentation and primary sources before relying on it.
What it is#
MCMC (Markov Chain Monte Carlo) is a family of algorithms for drawing samples from complex probability distributions — above all the posterior in Bayesian inference, which usually has no closed form. Rather than computing the distribution analytically, MCMC generates a long sequence of samples that approximate it: it lets you approximate posteriors by simulation when the math is intractable.
The two ideas in the name#
Markov chain — a sequence of states where the next depends only on the current one (not the full history); this structure is what lets the sampler walk through the distribution.
Monte Carlo — approximating integrals and expectations by repeated random sampling (the classic toy is estimating \(\pi\) from random points in a square).
Put together: build a Markov chain whose stationary distribution is the posterior, run it long enough, and its samples approximate that posterior.
The estimator#
Most Bayesian quantities are posterior expectations,
intractable in high dimensions — so MCMC approximates it by the sample average over the chain,
The algorithms#
Metropolis–Hastings — propose a move, then accept or reject it by a ratio of posterior densities.
Gibbs sampling — a special case that updates each parameter from its full conditional; natural for hierarchical models.
Hamiltonian Monte Carlo (HMC) — uses gradients to propose distant, high-acceptance moves; the engine of Stan.
NUTS (No-U-Turn Sampler) — an adaptive HMC that tunes step size and trajectory length automatically; the default in PyMC and Stan.
Example#
Coin toss with 7 heads, 3 tails and a \(\text{Beta}(2,2)\) prior gives a \(\text{Beta}(9,5)\) posterior. Here it’s closed-form, but if it weren’t, MCMC would draw thousands of samples of \(p\) to estimate the posterior mean, variance and credible interval.
Using it well#
MCMC handles very complex, high-dimensional posteriors and yields full samples (any posterior summary is then just a function of them). The costs: it is computationally expensive, needs tuning (burn-in, thinning, proposal scale), and requires convergence diagnostics — discard a burn-in, then check the potential-scale-reduction \(\hat{R}\) (near 1.0) and the effective sample size to confirm the chains have mixed. Versus variational inference, MCMC is asymptotically exact but slower; VI is faster but approximate.
Theme: Bayesian Inference · All terminology
Hint
Mind map — connected ideas
Variational Inference (VI) · Bayesian Inference. · Posterior · Beta Distribution · Bayesian Neural Networks (BNNs)
Hint
More in Bayesian Inference
Bayes’ Theorem · Bayesian Correction · Bayesian Decision Theory (BDT) · Bayesian Inference. · Bayesian Neural Networks (BNNs) · Binomial Likelihood · Gaussian Processes (GPs) · Marginal Likelihood (also called The Model Evidence or Integrated Likelihood) · Parameter(s) of Interest · Posterior · Posterior belief · Posterior Probability · Posterior probability of uplift · Prior Belief (or Prior Probability)
See also
Source article Adapted (context, re-expressed) in our own words from: MCMC (Markov Chain Monte Carlo) (insightful-data-lab.com).