Prediction asks about the world as observed
A supervised model estimates patterns such as “given what was observed, which outcome is likely?” A policy asks what would happen if an action changed while relevant conditions were comparable. That second question requires a counterfactual ordinary prediction does not identify.
A common cause can create a false lever
Support tickets and churn may both rise because a serious product failure caused them. Suppressing tickets does not remove the failure. Geography or device may proxy for access and policy differences that are predictive but inappropriate to manipulate.
An outcome can leave traces before it is recorded
Impending churn may reduce logins before cancellation, creating reverse causation. Forcing logins does not manufacture engagement. Selection also matters: a model trained only on approved applicants or treated patients cannot directly describe the outcomes of excluded people.
Draw the assumptions in a causal graph
Define population, action, comparison, outcome, and time horizon, then map plausible confounders, mediators, and colliders with domain experts. The diagram may reveal missing variables, forbidden controls, or an action that cannot ethically be randomized.
Prefer a randomized experiment when feasible
Random assignment can identify effects when interference, noncompliance, attrition, and measurement are handled correctly. Predeclare the outcome and analysis, monitor harms, and evaluate the policy rather than treating randomization as a magic label.
Observational methods require stronger assumptions
Matching, weighting, difference-in-differences, regression discontinuity, and instrumental variables solve different designs. Each depends on assumptions that must be stated and tested where possible; a sophisticated estimator cannot recover a missing identification strategy.
Uplift asks who responds differently
Uplift and heterogeneous-treatment models estimate variation in treatment effect, not merely outcome probability. They still require randomized or credibly causal data. High outcome risk does not necessarily identify the people whose outcome an intervention can change.
Evaluate prediction and policy separately
A predictive model can forecast demand, prioritize review, or define experimental strata. Measure who receives or misses the intervention and the resulting outcome rather than assuming higher AUC means better policy. Plan for feedback loops and collect post-intervention evidence; feature importance cannot provide the missing counterfactual.