A Ruler Is Not a Judge
firebrand diff aligns two sessions by step: request hashes, calls, tokens, timings, stop reasons. firebrand stats rolls up steps, calls, tokens, cuts, wall clock, tool errors. Those two are the rulers. There is no firebrand bench verb. None of these emit a quality score. Lifecycle facts only. A number that pretends to rank souls is the thing this family already refused on the workshop side — repair is reader-triggered, not scanner-ranked.
Blind comparison is N seats, N profiles, one brief. Judges, when they exist, sit in the instrument, not in this CLI. Firebrand-versus-Firebrand is still two notebooks, not a trophy.
What Replay Proves, Again
Replay proves the same recorded request normalizes the same way, or it proves it does not. It does not prove tomorrow's sample. A small-model cell that went 3/3 on one daemon release and 2/3 on another is two rows, both kept. Averaging them into 'the model is a B' is ranking by another door. Keep both rows. Name the daemon. Let a reader decide whether the work is done.