Ology

From Paper to Proof-Carrying Program

Methods for translating AI research papers into runnable, faithful, auditable code and responsibly maintaining it.

5 sessions · 15 readings · 2 views

Curated by Ben Lorenzo

The Paper Is Not the Specificationsession 1Making Ambiguity Executablesession 2Reproduction Is an Experimental Design Problemsession 3Evidence, Traces, and False Confidencesession 4From Generated Code to Scientific Infrastructuresession 5

Opening

Turning a research paper into runnable code sounds like an unusually tractable use of AI: give an agent the paper, let it write the implementation, run the experiments, and compare the results. Recent systems make parts of this workflow surprisingly plausible. PaperBench, for example, treats research replication as a structured engineering task rather than a vague coding exercise, while newer scientific-computing workflows show agents modernizing legacy research software and testing numerical implementations.

But a paper is not a complete specification. Important decisions may be distributed across prose, equations, figures, appendices, released code, undocumented conventions, and researcher intuition. An agent can produce code that runs while silently resolving ambiguity in the wrong direction. It can also optimize for a visible metric, misread an experimental setup, or express confidence despite lacking a valid scientific oracle.

The central difficulty is therefore not code generation alone. It is building a chain of evidence from claim to specification, from specification to implementation, from implementation to experiment, and from experiment to an auditable conclusion. The readings disagree, implicitly or explicitly, about how much structure should come from agents, how much should come from harnesses and tests, and where expert judgment remains irreducible.

This seminar follows that chain: how papers become executable plans, how ambiguity is managed during implementation, how reproduction becomes evidence rather than theater, how uncertainty is measured, and what responsible stewardship looks like after the first successful run.

Five questions worth arguing about

  1. 1Can an agent extract a paper’s real contribution without turning underspecified prose into unjustified assumptions?
  2. 2Should faithful implementation optimize for the paper’s written method, released code, or experimentally observed behavior?
  3. 3When does a passing reproduction demonstrate scientific validity rather than merely successful software execution?
  4. 4Can traces, rubrics, and evaluation harnesses make research code genuinely auditable?
  5. 5Who owns the maintenance, interpretation, and consequences of agent-generated scientific software?

The Sessions

1Session 1start here

The Paper Is Not the Specification

The problem

A research paper presents a compressed argument, not an executable contract. Its contribution may depend on details that are implicit, scattered across sections, or absent altogether. The first session asks whether agents can recover a usable specification without confusing interpretation with invention.

OpenAI’s PaperBench makes this problem concrete by decomposing replication into thousands of gradable outcomes, with rubrics developed alongside original paper authors. Its result is sobering: even strong agents struggle when they must understand the contribution, construct a codebase, and execute experiments from scratch. The benchmark treats faithful replication as a hierarchy of claims rather than a single pass/fail outcome.

Anthropic’s multi-agent research system approaches specification recovery through decomposition, parallel investigation, source evaluation, and citation-oriented synthesis. Hugging Face’s work on structured CodeAgents offers a different angle: if the agent’s actions are expressed as executable code, it gains flexibility and compositionality, but structure and parsing constraints become part of the reliability problem. Together, these readings raise a foundational question: does better representation solve ambiguity, or merely make assumptions easier to operationalize?

The next session moves from interpreting the paper to deciding what an implementation should actually do when the paper leaves room for judgment.

2Session 2Making Ambiguity Executable3 readings3Session 3Reproduction Is an Experimental Design Problem3 readings4Session 4Evidence, Traces, and False Confidence3 readings5Session 5From Generated Code to Scientific Infrastructure3 readings
From Paper to Proof-Carrying Program · Ology