Ology

Code Is Not a File: Reading Graphs of Software

Code graphs model source code through syntax, control flow, data flow, dependencies, and semantics for analysis, navigation, security, and tooling

5 sessions · 15 readings · 21 views

Curated by Rajeev Chhajer

The Shape of a Programsession 1Precision Has a Pricesession 2Findings Are Paths, Not Patternssession 3Search Is a Graph Problemsession 4Can Agents Query the Graph?session 5

Opening

Source code is text, but the questions developers ask about it rarely are. “Where is this function used?”, “What breaks if I change this type?”, “Can untrusted input reach this database call?”, and “Which files should an AI agent read?” all require relationships among syntax, symbols, control flow, data flow, dependencies, history, and usage.

The important recent change is not simply that AI systems write more code. It is that code intelligence has become part of the infrastructure around those systems. Code navigation indexes, semantic search engines, static-analysis databases, and repository-aware coding assistants all attempt to turn sprawling codebases into queryable structures—though they disagree about which relationships matter and how much precision is worth paying for.

A graph is therefore not one thing. A compiler-accurate symbol index, a control-flow graph, a taint-tracking database, a dependency graph, and a PageRank-style usage graph answer different questions. Some are fast but shallow; others are precise but difficult to build, update, or generalize across languages. Even “semantic understanding” often means a carefully selected approximation.

This seminar follows those choices from representation to query, from query to security finding, from search to developer tooling, and finally to AI agents that must navigate code rather than merely autocomplete it. The central question is: what kind of graph makes software legible enough for both analysis and action?

Five questions worth arguing about

  1. 1Should code intelligence optimize for one universal graph, or for many specialized representations with explicit boundaries?
  2. 2When does compiler-level precision justify the cost of building and maintaining a semantic code index?
  3. 3Can vulnerability detection remain useful without choosing between broad coverage and trustworthy paths?
  4. 4Does graph structure improve code search because it understands meaning, or because it supplies better ranking signals?
  5. 5Will AI coding agents benefit most from richer graphs, or from repositories designed to make their important relationships obvious?

The Sessions

1Session 1start here

The Shape of a Program

The problem

“Code graph” sounds like a single representation, but the readings expose a more difficult design problem: which aspects of a program deserve first-class structure? Files and symbols are useful for navigation; syntax trees support structural matching; compiler-derived relations support definitions and references; generalized intermediate representations promise portability across languages.

Beyang Liu’s account of LSP frames code intelligence as a systems-boundary problem. Separating language analysis from editors turns an otherwise multiplicative integration problem into a more manageable one, but the protocol does not eliminate the underlying question of what information a language server should expose. GitHub’s CodeGen approaches the issue from another direction: a scalable language-support pipeline needs representations that can normalize diverse languages without erasing their important differences.

SCIP then makes the trade-off operational. Search-based navigation can be available immediately but imperfect, while precise navigation can be compiler-accurate and cross-repository but requires language-specific indexing infrastructure. The session begins here because every later use—querying, security analysis, ranking, or AI retrieval—depends on what the representation preserves and what it discards.

2Session 2Precision Has a Price3 readings3Session 3Findings Are Paths, Not Patterns3 readings4Session 4Search Is a Graph Problem3 readings5Session 5Can Agents Query the Graph?3 readings