"Every reading points at the exact words it stands on. Change the rubric, and history is re-read — never rewritten."
Evidence Is the Anchor
The single most important field in a reading is evidence: the verbatim span, in its original language, that the reading rests on. "Mild" is a derived label; "a little stiff" is the truth it was derived from, and the reading carries the truth right beside the label. This is what keeps the typed layer honest. Any later analysis that cites a symptom must be able to trace back through the reading's evidence to Dad's actual words on an actual day — a claim you cannot trace to a span is a claim the engine is not allowed to make.
Normalization Without Erasure
Severity is normalized onto a small ladder — mild, moderate, severe, unspecified — and trend onto improving, worsening, stable, unspecified. But normalization happens on top of the evidence, never in place of it. The canonical field values are English so the aggregates in Track 6 can group them, while the evidence stays in whatever language Dad wrote. So "꽤 아팠다" becomes severity: moderate with the Korean span preserved. The comparable label and the citable truth coexist; neither is sacrificed to the other.
The Boundary Lives Inside the Rubric
One more thing the rubric carries: the no-diagnosis boundary, stated inside the rubric itself. The instruction that reads a crumb also tells the model, in the same breath, that it is producing observations and never a diagnosis. The safety line from Track 2 is not bolted on after the reading — it is part of the contract the reading is produced under. That is the pattern to copy: when a rule must always hold, put it where the work happens, not in a checker that runs afterward and hopes to catch a miss.