Esc
Quick summary
- Published benchmark scores have become a marketing surface, and the numbers rarely travel well between labs.
- When two teams report different results for the same evaluation, the cause is almost never the model weights.Where the variance comes fromPrompt formatting, sampling temperature, how many attempts are allowed and whether partial credit is given all move the final figure by several points.
- None of these are usually disclosed alongside the headline number.What a useful report looks likeThe evaluations worth trusting publish their harness, the exact prompts and the raw per-item outputs.

Comments 0
Join the discussion below