The Measure That Landed
Once seventy-eight items had reader scores, the same exercise that had failed all day produced something that worked. The winning measure came closer to the defect than any proxy had — a direct count rather than a stand-in: the share of the prose, in the language that is supposed to be the target, that is written in the source language's alphabet. Closer, not identical, and this lesson's own cautions say why: two items a reader scored 20/20 sit at 26 and 37 percent, where that alphabet is carrying identifiers and commands that are correct exactly as they stand.
That is not clever. It is a direct count of the thing the readers were describing when they said the worst items had left the meaning-carrying words alone and translated only the grammar. Correlation with reader totals came out strongly negative, and sixty-three percent of the bottom band was flagged against twenty-three percent of the top.
The Shape Is a Step, and the Shape Is the Whole Story
The correlation coefficient hides what matters. Binned, the relationship is not a slope: below roughly forty percent, items score well; past that line they collapse; and beyond it, more of the measured quantity predicts nothing further. Seventy percent is not worse than forty-five.
So the honest statement is a threshold, not a ranking — and that is not a stylistic preference, it is what the data supports. Using this to order a worklist would be reading structure into a region where the measure demonstrably has none. What orders the worklist is the reader's total. What the measure does is say which side of a line something falls on.
Three Cautions That Are Each Load-Bearing
It cannot be a gate. Good technical material runs high on this measure legitimately, because identifiers and commands are correct in the source language. Two items scoring perfectly with a reader sit at 26% and 37% — under the line, so the line as drawn passes them, and that is exactly how narrow the margin is. Tighten the threshold by three points to catch more of the bad band and you convict a perfect item.
It misses. One item scored a flat zero with a reader at a low measured share — bad for an entirely different reason, one no counting instrument can see. Passing this measure is not clearance.
One tempting variant is a dud. A related pattern that felt like the same signal correlated at almost nothing, and it is recorded specifically so nobody rebuilds it: the pattern appears everywhere, including in items that read perfectly.