One model, four candidates, one shelf slot
Track two gave you the instrument: a dtype label is evidence, and bytes-per-parameter arithmetic tests it. This lesson runs the whole procedure as a decision, because that is where the trap actually bites — not in understanding the rule, but in choosing among real repos under time pressure.
The scenario: a well-known diffusion model, and four public candidates for your archive's single master slot.
- The maker's original release — org account, BF16, shards summing to ~17 GB for an 8.9B-parameter model, paper agrees on sizes.
- The maker's "refresh" — same org, a week later, one commit, same stated dtype, same total size, message: "re-uploaded for consistency".
- A community mirror — byte sizes identical to (1), third-party account, no other provenance.
- A "space-saving BF16" — a fan account, same dtype claim, ~6 GB.
Running the procedure
Candidate 4 fails arithmetic. 8.9B params at 2 bytes ≈ 17.8 GB of tensors; 6 GB is Q8-sized. The label is decorative; this is a quantization wearing the original's name. Eliminated — or acquired later, deliberately, as a labeled rung.
Candidate 3 is redundant. Consistent with (1) but one step from the source with a frozen update story. It corroborates (1)'s sizes and earns nothing else. Record as corroboration; do not acquire.
Candidate 2 is the trap. Same label, same size, fresh commit — and zero story about what changed. A no-diff-size refresh is precisely the shape of a quiet re-export: re-tokenized, subtly re-sharded, or upcast from a lossier internal master. Nothing is provably wrong; nothing is provably identical; and "provably identical" is what a master slot requires. The procedure's answer: investigate before promoting — the hub publishes an LFS digest per file per revision, so compare the two revisions' digests file by file: identical digests everywhere prove a byte-for-byte copy, and any differing file marks a re-export. Read the release notes while you are there. If the investigation cannot close, the archive keeps the revision it can vouch for.
Candidate 1 wins — not by default, but by having every leg at once: arithmetic consistent, chain of custody complete, corroboration independent. That conjunction is what "authoritative" means operationally.
The habit to take away
Notice the shape of the whole exercise: it never required loading a model or running inference. Reconnaissance (sizes, history, accounts), arithmetic (params × bytes), and reading (cards, papers) settled everything. The dtype trap is not defeated by expertise — it is defeated by a checklist applied in order, before bytes move. That is the review desk's job: cheap questions first, expensive trust last.