AMP

Lena Marsh ·

Published benchmark scores have become a marketing surface, and the numbers rarely travel well between labs. When two teams report different results for the same evaluation, the cause is almost never the model weights.

Where the variance comes from

Prompt formatting, sampling temperature, how many attempts are allowed and whether partial credit is given all move the final figure by several points. None of these are usually disclosed alongside the headline number.

What a useful report looks like

The evaluations worth trusting publish their harness, the exact prompts and the raw per-item outputs. That lets a third party reproduce the score rather than take it on faith.

Until that becomes standard, treat a single benchmark figure as a claim rather than a measurement.

View full article