The Paper Is Not the Specification
The problem
A research paper presents a compressed argument, not an executable contract. Its contribution may depend on details that are implicit, scattered across sections, or absent altogether. The first session asks whether agents can recover a usable specification without confusing interpretation with invention.
OpenAI’s PaperBench makes this problem concrete by decomposing replication into thousands of gradable outcomes, with rubrics developed alongside original paper authors. Its result is sobering: even strong agents struggle when they must understand the contribution, construct a codebase, and execute experiments from scratch. The benchmark treats faithful replication as a hierarchy of claims rather than a single pass/fail outcome.
Anthropic’s multi-agent research system approaches specification recovery through decomposition, parallel investigation, source evaluation, and citation-oriented synthesis. Hugging Face’s work on structured CodeAgents offers a different angle: if the agent’s actions are expressed as executable code, it gains flexibility and compositionality, but structure and parsing constraints become part of the reliability problem. Together, these readings raise a foundational question: does better representation solve ambiguity, or merely make assumptions easier to operationalize?
The next session moves from interpreting the paper to deciding what an implementation should actually do when the paper leaves room for judgment.
Readings