The Equation
Read aloud: "the probability of A given B equals the probability of B given A, times the probability of A, divided by the probability of B." That's Bayes' rule. It tells you how to update a belief about A when you observe B.
The Mental Model
- Prior — what you believed about A before seeing the evidence.
- Likelihood — how likely the evidence B is, assuming A were true.
- Posterior — your updated belief about A after seeing B.
- Evidence — a normalizer that ensures probabilities sum to 1.
The Classic Counter-Intuition Example
1% of the population has a disease. A test for the disease is 99% accurate (99% true positive rate, 99% true negative rate). You test positive. What's the probability you actually have it?
Most people guess 99%. Bayes says ~50%. The reason: the prior is 1% — diseases are rare — and the false-positive rate (1%) applied to the 99% of healthy people produces almost as many false positives as the 99% true positive rate produces true positives. Updating from a low prior takes more evidence than your gut estimates.
Bayes in ML
- Naive Bayes classifiers — apply Bayes directly with independence assumptions.
- Bayesian neural networks — treat weights as distributions, update with data.
- Variational inference — approximate intractable posteriors with simpler distributions.
- RLHF / DPO — modern LLM training updates a prior policy toward a posterior aligned with human preferences.