What Counts as Success?
The problem
The first mistake in building an LLM judge is treating evaluation as a neutral measurement step. Before choosing a prompt or score, designers must decide what claim the evaluation is meant to support: capability, safety, usefulness, reliability, or performance in a particular workflow. Different claims require different tasks, contexts, baselines, and standards of evidence.
Anthropic’s account of evaluation difficulty emphasizes that benchmark scores can fail to represent the behavior teams actually care about. Google DeepMind approaches the problem from a broader angle, arguing that capability evaluations alone miss what happens during human interaction and after systems become embedded in institutions. The disagreement is productive: one reading focuses on measurement validity inside the system, while the other asks whether the system boundary itself has been drawn too narrowly.
The shift toward agents makes the question more urgent. A response judged in isolation may look correct even when the system used the wrong tool, ignored a constraint, or produced a harmful downstream effect. Before constructing a judge, then, we need to decide what kind of object is being judged—and what evidence would justify the conclusion.
Readings