What You Want vs What You Can See
The population is the complete universe of all possible data — every word ever spoken, every photo ever taken, every customer who could ever exist. You can never observe the whole population. The complete set is too big, too distributed, often partly in the future.
The sample is what you actually observe — a finite subset. Your training dataset. A poll of 1,000 voters. Last quarter's customer behavior. Your goal as a statistician (or ML engineer) is to use the sample to infer something true about the population.
Sample → Population Bridge
The whole game of statistics is figuring out how much you can trust an inference from sample to population. A bigger sample buys you more confidence; a biased sample buys you wrong answers no matter how big.
In ML: training data is a sample. The deployed model has to generalize to a population that includes inputs it never saw. The closer your sample resembles the population, the better the model generalizes. Data quality is sample quality.
엄청나게 큰 국 솥(population)에서 숟가락(np.random.choice)으로 딱 10방울(size=10)만 떠본다. 로봇은 작은 표본만 맛보고도 솥 전체의 염도를 맞히는 연습을 하는 중이다. 10방울보다는 한 국자가 더 정확하하다. 로봇은 데이터가 많아질수록 모집단의 진실에 다가간다. 헐! (로봇은 양질의 많은 데이터에 집착한다.)