The Context Budget Is the Architecture
The problem
A larger context window appears to promise a simple solution: keep more history, tool output, and reference material available to the model. But the readings challenge that intuition from different directions. Anthropic defines context engineering as the continual curation of the tokens available at inference time, while Google argues that context should be treated as a compiled, per-call view over richer underlying state.
Google’s earlier account of Gemini’s long-context work supplies an important counterpoint. Larger windows genuinely enable new behaviors, including reasoning over large codebases, long videos, and specialized reference material. Yet the later engineering perspective suggests that raw capacity does not eliminate relevance, cost, latency, or “lost in the middle” problems.
OpenAI’s Codex experience makes the issue concrete: repository knowledge became more useful when organized as a navigable map rather than a giant instruction manual. The tension is not between short and long context in the abstract, but between context as accumulated history and context as deliberately structured evidence. That distinction leads directly to the next question: who decides what enters the window?