Ology

When the Judge Becomes the Benchmark

Fundamentals of using large language models as judges to evaluate AI-generated outputs, including rubrics, prompting, comparison, calibration, reliability, and bias

5 sessions · 12 readings · 40 views

Curated by Rajeev Chhajer

What Counts as Success?session 1Writing the Standardsession 2Scores, Rankings, and the Shape of Judgmentsession 3Can the Judge Be Trusted?session 4The Judge’s Blind Spotssession 5

Opening

Large language models are increasingly being asked to evaluate other models—not only by assigning a score, but by deciding whether an answer is useful, safe, accurate, aligned, or ready for deployment. This is attractive because open-ended outputs are difficult to grade with exact-match metrics, while human review is expensive, slow, and often inconsistent.

The practice has also changed as the systems being evaluated have changed. Recent evaluations increasingly involve realistic conversations, domain-specific rubrics, multi-turn agents, tool use, and deliverables judged against professional standards. In this setting, the judge is no longer merely checking an answer; it is interpreting a task, a context, and a theory of what success means.

That makes evaluation design inseparable from prompt design and measurement theory. A judge may prefer longer answers, favor one response position, reward stylistic fluency, or reproduce the assumptions embedded in its rubric. Even apparently impressive agreement with human ratings can conceal systematic blind spots, especially when the judge and the evaluated model share training data or behavioral tendencies.

The five sessions follow this problem from first principles to deployment: what exactly should be measured, how should judgments become rubrics, when are pairwise comparisons preferable to scores, how can judges be calibrated, and what remains fundamentally resistant to automation?

Five questions worth arguing about

  1. 1What makes an evaluation measure the system’s real-world behavior rather than merely its performance under a convenient test?
  2. 2Should rubrics encode expert judgment explicitly, or preserve room for a judge to recognize quality designers failed to anticipate?
  3. 3When do pairwise preferences reveal more than numerical scores, and when do they make evaluation less interpretable?
  4. 4How much human validation is enough before an LLM judge becomes trustworthy at scale?
  5. 5Can automated judges expose behaviors that humans miss without reproducing the biases and blind spots of the systems they evaluate?

The Sessions

1Session 1start here

What Counts as Success?

The problem

The first mistake in building an LLM judge is treating evaluation as a neutral measurement step. Before choosing a prompt or score, designers must decide what claim the evaluation is meant to support: capability, safety, usefulness, reliability, or performance in a particular workflow. Different claims require different tasks, contexts, baselines, and standards of evidence.

Anthropic’s account of evaluation difficulty emphasizes that benchmark scores can fail to represent the behavior teams actually care about. Google DeepMind approaches the problem from a broader angle, arguing that capability evaluations alone miss what happens during human interaction and after systems become embedded in institutions. The disagreement is productive: one reading focuses on measurement validity inside the system, while the other asks whether the system boundary itself has been drawn too narrowly.

The shift toward agents makes the question more urgent. A response judged in isolation may look correct even when the system used the wrong tool, ignored a constraint, or produced a harmful downstream effect. Before constructing a judge, then, we need to decide what kind of object is being judged—and what evidence would justify the conclusion.

2Session 2Writing the Standard3 readings3Session 3Scores, Rankings, and the Shape of Judgment3 readings4Session 4Can the Judge Be Trusted?2 readings5Session 5The Judge’s Blind Spots2 readings