The Magic Trick
Take any distribution — uniform, exponential, weird custom thing. Sample from it many times. Compute the mean of each sample. The distribution of those sample means, regardless of the original distribution, will approach a normal distribution as your sample size grows.
Read that again. Regardless of the original distribution. That's the Central Limit Theorem. It's why the bell curve shows up in places that have nothing to do with bells.
Why Heights Are Bell-Shaped
Adult height is influenced by hundreds of genetic factors plus thousands of environmental ones. Each individually is a small random contribution. The CLT says: when you sum or average many independent random factors, the result tends toward a normal distribution. Heights are normal because they're the sum of many small factors. Same for measurement errors. Same for IQ. Same for sums of dice rolls.
Implications for ML
- Why we assume Gaussian noise everywhere. Sensor noise, model residuals, batch averages — they're often sums of many tiny effects, so the CLT delivers a Gaussian-shaped result.
- Why batch normalization works. Averaging activations over a batch invokes the CLT — the batch mean has a Gaussian-ish distribution.
- Why confidence intervals on means are normal-based. Even if individual data is non-normal, the mean of enough samples is normal-ish.