The Score Is Not the Claim
The problem
The first mistake in evaluation is treating a numerical result as if it had a self-evident meaning. Anthropic’s critique of common evaluation practice shows how apparently simple benchmarks can be distorted by contamination, formatting choices, inconsistent prompting, mislabeled questions, and unanswerable items. The score may be accurate as a measurement of the test procedure while still failing to support the broader capability claim.
Hugging Face approaches the same problem from a wider methodological angle, separating automated benchmarks, human judgments, and model-based judges. Its central caution is that “general capability” is often an ill-defined construct, while practical evaluation can still provide useful signal for regression testing and task-specific comparison. Anthropic’s statistical treatment adds another layer: even a well-designed benchmark difference may reflect sampling luck unless uncertainty and experimental design are reported properly.
Together, these readings establish the seminar’s baseline distinction between a result and the proposition it is supposed to justify. The next question is what happens when the evaluation environment itself changes or becomes visible to the system being tested.
Readings