Discrete Bayesian Examples – Genetics and Spell Checking (with θ)#
Part 1 · Stage 1 · 🎲 The Bayesian Idea · Lesson 004 of 144 · beginner
◀ Previous · Bayesian Inference · Next · Probability as a Measure of Uncertainty ▶ · ↑ Section
Important
✨ AI-generated content. This page was written with the assistance of an AI language model and is provided as a learning aid. Despite careful review, it may still contain mistakes, omissions, or out-of-date information. Whether you are new to the topic, a team lead, or a senior practitioner, treat it as a starting point rather than an authoritative reference: read it critically and independently verify anything you act on (code, commands, figures, and factual claims) against official documentation and primary sources before relying on it.
Bayes’ rule on a discrete unknown#
The clearest way to see Bayes’ rule work is when \(\theta\) takes only a few discrete values. No integrals, no sampling — just arithmetic on probabilities. Two classic examples, both in Gelman’s first chapter, show the whole logic of updating.
Genetics: the carrier problem#
Hemophilia is X-linked recessive. A woman whose brother is affected has a carrier status \(\theta \in \{0, 1\}\) with prior \(\Pr(\theta = 1) = \tfrac{1}{2}\). Suppose she then has two unaffected sons, \(y = (0, 0)\). A carrier passes the gene with probability \(\tfrac{1}{2}\) per son, so the likelihoods are \(p(y \mid \theta = 1) = (\tfrac{1}{2})^2 = \tfrac14\) and \(p(y \mid \theta = 0) = 1\). Bayes’ rule updates her carrier probability:
Two healthy sons drop the probability from 0.50 to 0.20 — evidence, not proof. Each further unaffected son multiplies the odds again; the update is sequential, and today’s posterior is tomorrow’s prior.
Spell checking: which word was meant?#
The same rule powers a spell checker. Seeing the typed string radom, the candidate corrections
\(\theta \in \{\text{random}, \text{radon}, \text{radom}\}\) are scored by
\(p(\theta \mid y) \propto p(y \mid \theta)\, p(\theta)\): the prior is how common each word is
in the language, and the likelihood is how likely that typo is given the intended word.
prior = {"random": 7.6e-5, "radon": 6.1e-6, "radom": 1.2e-7} # word frequencies
like = {"random": 0.00193, "radon": 0.000143, "radom": 0.975} # typo model
post = {w: prior[w] * like[w] for w in prior}
Z = sum(post.values())
{w: p / Z for w, p in post.items()} # normalised posterior over intended words
A rare word needs a much better typo-likelihood to win. Both examples make the same point: the prior supplies context, the likelihood supplies evidence, and the posterior is their compromise — the theme of the next stage.
Hint
Related lessons: Bayesian Inference · Probability as a Measure of Uncertainty · Posterior as a Compromise Between Data and Prior Information · Estimating a Probability from Binomial Data
See also
Source article Adapted (context, re-expressed) in our own words from: https://insightful-data-lab.com/2025/11/08/discrete-bayesian-examples-genetics-and-spell-checking-with-%ce%b8/ (insightful-data-lab.com).