"A σ-based statement on fat-tailed data is mathematically valid arithmetic about a fantasy world."
The Lens Has a Domain
The σ lens works beautifully where the bell-curve assumption holds. It works badly — and often misleadingly — where it doesn't. Financial returns, social reach, file sizes, response times, wealth, network traffic, and earthquake measures can be skewed or heavy-tailed, but each distribution must be checked for a named variable and dataset.
For variables with finite variance, computing σ can be valid even when the distribution is nonnormal; with infinite variance or unstable estimates, σ itself may not be well-defined or reliable. The interpretation 'this happens 0.13% of the time' is a fantasy borrowed from a normal distribution that doesn't fit the data. The arithmetic is fine; the conclusion is wrong.
How the Mismatch Hides
A limited sample can hide tail behavior, and a calm period may underrepresent stress regimes. A histogram's center alone cannot validate a tail model. The σ computed from it is well-defined. The 95% interval computed from that σ covers most days. Everything seems fine — until a crisis day arrives, at which point the 'rare' event happens, and another arrives the next day, and another the day after that. The σ lens was sampled from the calm middle of a fat-tailed distribution; it has nothing to say about the tails it didn't see.
This is what makes the failure mode insidious. The σ lens reports zero red flags on quiet data. It only fails when it matters most: in the tails, where the consequential events live, and where its calibration is most divorced from reality.
The Citizen's Two-Question Test
Before applying the σ lens to any new dataset:
- What is the underlying process? What process, dependence, mixture, and tail behavior generated the data? A sum-of-small-effects story alone does not establish a normal marginal distribution.
- Have you seen the tails? A calm sample of fat-tailed data looks normal. The lens calibrated on it will lie when the tail event arrives. Stress-test your σ assumption by asking: if a 5σ event happened tomorrow, would you genuinely consider it a 1-in-3.5-million coincidence, or would you suspect the model?