Exchangeability and hierarchical models#
Part 1 · Stage 5 · 🏛️ Hierarchical Models · Lesson 034 of 144 · beginner
◀ Previous · Constructing a Parameterized Prior Distribution · Next · Bayesian analysis of conjugate hierarchical models ▶ · ↑ Section
Important
✨ AI-generated content. This page was written with the assistance of an AI language model and is provided as a learning aid. Despite careful review, it may still contain mistakes, omissions, or out-of-date information. Whether you are new to the topic, a team lead, or a senior practitioner, treat it as a starting point rather than an authoritative reference: read it critically and independently verify anything you act on (code, commands, figures, and factual claims) against official documentation and primary sources before relying on it.
The licence to pool#
Hierarchical models let groups borrow strength from one another. What entitles them to? The answer is exchangeability: the assumption that the group parameters are, before seeing the data, interchangeable — that nothing in the labelling distinguishes them.
The definition#
Parameters \(\theta_1, \dots, \theta_J\) are exchangeable if their joint distribution is invariant under any permutation \(\pi\) of the indices:
It is a statement of ignorance, not of equality. The schools may differ enormously; you simply have no prior reason, before the data, to expect school 3 to differ from school 7 in any particular direction. Independent and identically distributed is a special case; exchangeability is weaker and more useful.
De Finetti’s bridge#
De Finetti’s theorem makes the connection precise: an exchangeable sequence can (in the infinite case) be represented as iid draws given some parameter, mixed over a distribution for that parameter,
Read the right-hand side: that is the hierarchical model of the last lesson. Exchangeability does not merely permit hierarchy — it implies it. The population distribution \(p(\theta \mid \phi)\) is not an extra assumption bolted on; it is what exchangeability means.
Three pooling regimes#
The hyperparameter governing spread (call it \(\tau\)) interpolates between two extremes, and the data choose where to sit:
\(\tau \to 0\) — complete pooling: all \(\theta_j\) collapse to a common value; group identity is ignored.
\(\tau \to \infty\) — no pooling: each group is estimated alone, noisily.
\(0 < \tau < \infty\) — partial pooling: each estimate is shrunk toward the population mean by an amount reflecting its own precision and the between-group spread.
When exchangeability fails#
If you do know something distinguishing — schools differ by funding, counties by population — then the parameters are not exchangeable as they stand. The repair is not to abandon the hierarchy but to make the known structure explicit: model \(\theta_j\) as depending on covariates \(x_j\), and assume exchangeability of the residuals. This is conditional exchangeability, and it is the road from hierarchical models to hierarchical regression in Part IV.
Hint
Related lessons: Constructing a Parameterized Prior Distribution · General Notation for Statistical Inference · Normal model with exchangeable parameters · Regression coefficients exchangeable in batches
See also
Source article Adapted (context, re-expressed) in our own words from: https://insightful-data-lab.com/2025/11/09/exchangeability-and-hierarchical-models/ (insightful-data-lab.com).