A Day of Plausible Ideas, All Wrong
Before any of the measurement in this track existed, a full day went into building a mechanical detector for the defect the first track described. Every proxy tried was reasonable. Every one failed, and the failures are recorded in a table nobody is allowed to delete.
Thin content. Flag items whose sections fall below a length floor. Result: it flagged sixty-nine of the seventy-seven items that existed in the fragment-first format the scan could read and thousands of sections. It paints the corpus and ranks nothing.
A punctuation marker. Look for a specific construction that felt like a signature of the defect. Result: sixty-six of seventy-seven items, over fifteen hundred hits, and the hits were dominated by entirely legitimate uses.
A lexicon-free non-word detector. Clever: find words used in two grammatical patterns that only co-occur in malformed constructions, then filter by rarity. Result: thirty candidates, one true positive. Three percent, and the false positives were ordinary words.
An existing screening scan, reused as a ranker. Result: no correlation at all. Items scoring high on the scan averaged fewer real findings than items that had already been cleaned.
Why the Table Has to Be Kept
Not as an obituary. Every one of those ideas is what a competent person reaches for on day one, which means that without the record the next person spends the same day rebuilding one of them and arriving at the same number.
So each row has to carry what it measured, what the result was, and — this is the part people leave out — why it failed, in a sentence. "No correlation" is a result. "Dominated by legitimate uses, because the construction is also how the language forms ordinary compounds" is a reason, and a reason transfers to the next idea in a way a number does not.
The Line to Read the Table With
One important caveat, and it is easy to get backwards. This table is not evidence that mechanical detection is impossible. It is a record of proxies tried with nothing to validate against. Every one of them was a guess about what the defect looks like, evaluated by whether its output felt right.
The moment ground truth existed — a set of items scored by a reader — the same exercise produced something that landed. Which is the actual lesson: the proxies did not fail because the domain resists measurement. They failed because they were built before anything could tell them they were wrong.