"The data you analyzed is not a random sample of the population you cared about. Almost no data is. Knowing the filter is the citizen-statistician's primary defense."
What Selection Bias Is
Selection bias occurs when the process that produced your sample is correlated with the outcome you are trying to measure. The classic forms:
- Self-selection: subjects volunteer to participate. Volunteers differ from non-volunteers in ways that often correlate with the outcome.
- Non-response bias: subjects refuse to answer. The refusers differ from the responders.
- Survivorship: only some subjects are reachable for measurement (already covered in lesson 3).
- Geographic / demographic filtering: the sample only includes people from places or groups that are accessible to the researcher.
- Convenience sampling: the researcher uses 'whoever is around,' which is rarely a representative slice.
Each of these produces a sample whose composition is no longer the population's, and inferences drawn from the sample do not generalize.
The 1948 Election Polling Disaster
In the 1948 US presidential election, every major pre-election poll predicted Dewey would defeat Truman. The polls were spectacularly wrong. The cause was selection bias: polls were conducted by telephone, and in 1948, telephones were disproportionately held by wealthier, more Republican voters. The poll sample was systematically skewed toward Republican-leaning respondents, and the projection from the sample to the population was therefore biased. The 'Dewey Defeats Truman' headline became one of the most famous photographs in journalism history, and the methodology of political polling changed substantially as a result.
The Modern Form: Online Polls and 'Engagement'
Modern online polls and social-media 'sentiment' metrics suffer the same selection bias. People who fill out online polls are not a random slice of the population; they are more engaged, more opinionated, more online. Polls that draw from one platform (Twitter/X versus Reddit versus Instagram) draw from different demographic and ideological distributions. Aggregating across platforms or weighting carefully can partially correct, but the underlying filter never fully disappears. 'Internet sentiment' is the sentiment of the people who chose to express it, which is not the same as the sentiment of the population.