The Equation
Bayes' rule relates the probability of a hypothesis after evidence to the likelihood of that evidence under the hypothesis, the prior probability, and a normalizing evidence term.
The Mental Model
- Prior : the probability assigned before observing .
- Likelihood : how probable the observed evidence is if holds.
- Posterior : the updated probability after observing .
- Evidence : the normalizer obtained across the possible causes of .
A Counterintuitive Test Result
Suppose 1% of a population has a disease, and a test has 99% sensitivity and 99% specificity. A positive result does not imply a 99% chance of disease. Among 10,000 people, about 99 of the 100 people with the disease test positive, while about 99 of the 9,900 healthy people produce false positives. That leaves about 198 positive results split nearly evenly between true and false positives, yielding a posterior near 50% under these assumptions.
Bayesian and Non-Bayesian ML Uses
- Naive Bayes classifiers apply Bayes with conditional-independence assumptions.
- Bayesian neural networks place distributions over weights or functions and update them with data.
- Variational inference approximates difficult posterior distributions with tractable families.
- Preference training methods such as RLHF and DPO update a policy using preference signals, but the resulting policy is not automatically a Bayesian posterior.
Bayes' rule is an exact relationship between prior, likelihood, evidence, and posterior. Calling any update a “Bayesian update” requires those probabilistic pieces, not just a before-and-after model.