Ology

Beyond the Score: Evaluating AI Systems That Act

Methods for evaluating AI systems and agents using benchmarks, tool-use tests, reliability and safety measures, human judgments, and real-world deployment

5 sessions · 17 readings · 15 views

Curated by Rajeev Chhajer

Scores Under Suspicionsession 1The Agent Is the Whole Experimentsession 2Can the Judge Be Trusted?session 3From Tasks to Worldssession 4Evals as Operating Infrastructuresession 5

Opening

AI evaluation used to look deceptively simple: pose a question, compare the answer with a reference, and report a score. That picture is breaking down. Modern systems browse, call tools, edit files, maintain state, recover from errors, and sometimes act inside consequential workflows. Their performance depends not only on the model, but also on the harness, environment, budget, evaluator, and opportunities available during the test.

Recent evaluation work makes the problem more concrete. OpenAI’s analysis of coding benchmarks finds that a substantial fraction of tasks are broken or ambiguously scored. Google DeepMind’s double-blind evaluation work treats contamination as a security and infrastructure problem, not merely a dataset-quality issue. Meanwhile, agent benchmarks increasingly measure trajectories, cost, tool use, and environmental state rather than final answers alone.

The evaluator is also part of the system under study. Human experts provide valuable judgment but are expensive and inconsistent; model judges scale more easily but inherit biases, miss subtle failures, and can agree with non-experts more readily than specialists. The practical response has been neither to abandon automation nor to trust it blindly, but to build layered evaluation pipelines that expose uncertainty and preserve human oversight where it matters.

This seminar follows that shift from scores to systems: what benchmark results mean, how agents should be tested in environments, when evaluators deserve trust, how long-horizon behavior can be measured, and how evaluation becomes an operational discipline rather than a release-day ritual.

Five questions worth arguing about

  1. 1When a benchmark is broken, contaminated, or gameable, what exactly does its score still tell us?
  2. 2Should agent evaluations reward successful outcomes, faithful procedures, or the ability to recover from surprises?
  3. 3When human judgment is scarce and model judges are biased, what evidence justifies automating evaluation?
  4. 4Can short, repeatable tasks measure the reliability and safety of agents acting over long horizons?
  5. 5How should evaluation pipelines change decisions when production failures, costs, and benchmark scores disagree?

The Sessions

1Session 1start here

Scores Under Suspicion

The problem

A benchmark score appears objective because it compresses many trials into one number. But the compression can conceal whether the task was well specified, whether the tests measured the intended capability, and whether the model solved the problem rather than exploited an artifact. The first session therefore asks us to treat validity as an empirical property of an evaluation, not as something guaranteed by publication or popularity.

Hugging Face’s overview frames the familiar tradeoff among automated benchmarks, human evaluation, and model judges, while emphasizing contamination and evaluator bias. OpenAI’s coding-evaluation post supplies a sharper case study: human reviewers and investigator agents uncovered broken tasks involving underspecified prompts, overly strict tests, and low test coverage. The lesson is not simply that benchmarks fail, but that increasingly capable models make it possible—and necessary—to audit the benchmark itself.

Google DeepMind approaches contamination as an information-security problem. Its double-blind evaluation design keeps both proprietary models and evaluation prompts hidden inside a cryptographically protected environment. OpenAI’s GDPval offers a contrasting response: move closer to economically valuable, expert-authored work and compare outputs with professional graders. Together, these readings set up the central tension for the seminar: should evaluation become more secure, more realistic, or both?

2Session 2The Agent Is the Whole Experiment3 readings3Session 3Can the Judge Be Trusted?3 readings4Session 4From Tasks to Worlds3 readings5Session 5Evals as Operating Infrastructure4 readings