Scores Under Suspicion
The problem
A benchmark score appears objective because it compresses many trials into one number. But the compression can conceal whether the task was well specified, whether the tests measured the intended capability, and whether the model solved the problem rather than exploited an artifact. The first session therefore asks us to treat validity as an empirical property of an evaluation, not as something guaranteed by publication or popularity.
Hugging Face’s overview frames the familiar tradeoff among automated benchmarks, human evaluation, and model judges, while emphasizing contamination and evaluator bias. OpenAI’s coding-evaluation post supplies a sharper case study: human reviewers and investigator agents uncovered broken tasks involving underspecified prompts, overly strict tests, and low test coverage. The lesson is not simply that benchmarks fail, but that increasingly capable models make it possible—and necessary—to audit the benchmark itself.
Google DeepMind approaches contamination as an information-security problem. Its double-blind evaluation design keeps both proprietary models and evaluation prompts hidden inside a cryptographically protected environment. OpenAI’s GDPval offers a contrasting response: move closer to economically valuable, expert-authored work and compare outputs with professional graders. Together, these readings set up the central tension for the seminar: should evaluation become more secure, more realistic, or both?