Ology

When the Score Lies: A Seminar on Reliable AI Evaluation

Measuring AI evaluation reliability, validity, drift, and disagreement to calibrate known, uncertain, and unsupported claims

5 sessions · 15 readings · 1 view

Curated by Ben Lorenzo

The Score Is Not the Claimsession 1When the Test Changes the Subjectsession 2Who Gets to Be the Judge?session 3Disagreement Is a Measurementsession 4Rewarding “I Don’t Know”session 5

Opening

AI evaluation is no longer just a matter of running a benchmark and reporting a percentage. Modern systems use tools, multi-step reasoning, simulated environments, automated graders, human reviewers, and dynamically generated tasks. Each layer can change what the evaluation actually measures.

Recent work has also made the distinction between competence and epistemic discipline harder to ignore. A model that answers more questions may be less reliable if it guesses when it should abstain. A model-based grader may scale evaluation dramatically while introducing its own biases, blind spots, and changing standards. Even expert judgments can disagree because the task, rubric, or evidence is underspecified.

The result is a measurement problem with several interacting failure modes: contamination, broken tasks, reward hacking, judge disagreement, distribution shift, and incentives that reward confident error over honest uncertainty. Reliability therefore requires more than agreement with a reference answer; it requires knowing what claim an evaluation supports, how stable that claim is, and when the evaluator itself should not be trusted.

This seminar follows that chain from benchmark scores to trustworthy claims: what makes an evaluation valid, how validity decays, when automated judgments deserve confidence, what disagreement reveals, and how evaluation can teach systems to distinguish knowledge from uncertainty and unsupported invention.

Five questions worth arguing about

  1. 1When does a benchmark score become evidence about a capability rather than evidence about a particular test-taking procedure?
  2. 2Can an evaluation remain valid when models adapt to its prompts, harness, scorer, or deployment environment?
  3. 3What evidence should justify replacing expert judgment with an automated evaluator?
  4. 4Should disagreement between evaluators be treated as noise, or as information about the claim being measured?
  5. 5Can evaluation metrics reward appropriate uncertainty without making systems evasive or unusably cautious?

The Sessions

1Session 1start here

The Score Is Not the Claim

The problem

The first mistake in evaluation is treating a numerical result as if it had a self-evident meaning. Anthropic’s critique of common evaluation practice shows how apparently simple benchmarks can be distorted by contamination, formatting choices, inconsistent prompting, mislabeled questions, and unanswerable items. The score may be accurate as a measurement of the test procedure while still failing to support the broader capability claim.

Hugging Face approaches the same problem from a wider methodological angle, separating automated benchmarks, human judgments, and model-based judges. Its central caution is that “general capability” is often an ill-defined construct, while practical evaluation can still provide useful signal for regression testing and task-specific comparison. Anthropic’s statistical treatment adds another layer: even a well-designed benchmark difference may reflect sampling luck unless uncertainty and experimental design are reported properly.

Together, these readings establish the seminar’s baseline distinction between a result and the proposition it is supposed to justify. The next question is what happens when the evaluation environment itself changes or becomes visible to the system being tested.

2Session 2When the Test Changes the Subject3 readings3Session 3Who Gets to Be the Judge?3 readings4Session 4Disagreement Is a Measurement3 readings5Session 5Rewarding “I Don’t Know”3 readings