Ology

What Does It Mean to Understand Code with AI?

Evaluating how AI coding tools affect code comprehension, reasoning, explanation, modification, and retained independent understanding.

5 sessions · 15 readings · 16 views

Curated by Ryan Lingo

From Plausible Output to Demonstrated Understandingsession 1The Metric Is Not the Outcomesession 2Realism, Control, and the Cost of Clean Evidencesession 3The Tool Leaves the Roomsession 4Designing the Study That Survives the Tool’s Evolutionsession 5

Opening

AI coding tools have moved from autocomplete toward explanation, debugging, refactoring, codebase navigation, and increasingly autonomous implementation. That shift changes the question from whether a tool can produce working code to whether the human using it can still inspect, explain, challenge, and extend what was produced.

The distinction matters because speed, output volume, satisfaction, and understanding can move in different directions. Developers may complete a task faster while becoming reviewers of unfamiliar output, or report large productivity gains while controlled experiments find slower completion. A tool can also make previously uneconomical work possible without making the user more capable of reasoning about the resulting system.

Recent studies have begun to separate these outcomes. Anthropic’s randomized study tests debugging, code reading, code writing, and conceptual knowledge; METR contrasts algorithmic benchmark scores with manual judgments and real task completion; Google’s learning experiments add delayed retention tests rather than relying only on immediate performance. Together, they suggest that “understanding” is not a single variable and that no single productivity metric can stand in for it.

A rigorous research program therefore needs operational definitions, behavioral measures, realistic controls, and tests conducted after assistance is removed. The five sessions investigate how to build such a program without confusing perceived usefulness, task completion, code correctness, and durable independent competence.

Five questions worth arguing about

  1. 1Which observable behaviors best distinguish understanding code from merely accepting plausible code?
  2. 2When do speed, output, satisfaction, and engineering value become misleading substitutes for one another?
  3. 3How much realism should an experiment sacrifice to obtain cleaner causal evidence?
  4. 4Can AI-assisted coding produce durable understanding, or only short-term task success?
  5. 5What study design could detect both skill formation and skill displacement as tools evolve?

The Sessions

1Session 1start here

From Plausible Output to Demonstrated Understanding

The problem

“Code understanding” is often treated as an intuitive concept, but a useful experiment must turn it into observable performance. Reading code, diagnosing a bug, explaining a design choice, writing an extension, and recognizing an inappropriate abstraction are related abilities, not interchangeable outcomes.

Anthropic’s randomized study provides one of the clearest operational decompositions currently available. Its assessment separates debugging, code reading, code writing, and conceptual questions, while its task uses an unfamiliar Python library so that participants must form a new mental model rather than rely only on prior expertise. The study’s central contribution is methodological: it treats understanding as a family of skills relevant to oversight, not simply as confidence or familiarity.

Google’s account of AI-assisted software engineering complicates this framework by observing that the developer increasingly becomes a reviewer. That may make review behavior itself a central outcome: Can a person identify subtle defects, explain why an implementation fits the surrounding system, and decide when not to trust a suggestion? GitHub’s earlier work adds another warning: even “productivity” requires a theory of what counts as meaningful developer performance.

The next question is whether these measures can be connected to outcomes that organizations actually care about, rather than remaining isolated laboratory scores.

2Session 2The Metric Is Not the Outcome3 readings3Session 3Realism, Control, and the Cost of Clean Evidence3 readings4Session 4The Tool Leaves the Room3 readings5Session 5Designing the Study That Survives the Tool’s Evolution3 readings