Ology

The Agent’s Second Mind

AI agent context management across conversations and workflows, including memory, retrieval, summarization, persistence, personalization, tool results, privacy, and reliability

5 sessions · 15 readings · 11 views

Curated by Rajeev Chhajer

The Context Budget Is the Architecturesession 1Retrieval or Anticipation?session 2What Should Survive?session 3Crossing the Windowsession 4Trusting the Residuesession 5

Opening

AI agents are no longer confined to short exchanges. They browse, call APIs, inspect files, revise plans, delegate subtasks, wait for approvals, and resume work hours or weeks later. The central engineering problem is therefore no longer how to write a good prompt, but how to decide what an agent should be allowed to remember, retrieve, ignore, verify, and carry forward.

Recent systems increasingly separate durable state from the context shown to a model at any particular step. Anthropic describes this as context engineering; Google frames context as a compiled view over sessions, memories, artifacts, and tools; OpenAI’s internal systems treat repository structure, tool outputs, institutional knowledge, and user corrections as distinct sources of grounding.

But “memory” is not one thing. A conversation transcript, a task state machine, a retrieved document, a tool result, a filesystem, and a user preference have different lifecycles and different failure modes. Compressing them into one growing prompt may preserve tokens while destroying provenance, freshness, or control.

The deeper question is whether context management should be understood as retrieval, state management, software architecture, or a form of cognitive infrastructure. These sessions follow that question from the attention budget of a single model call to the reliability and security of memories that survive across tasks.

Five questions worth arguing about

  1. 1Is better agent context mainly a problem of expanding windows, or of making fewer tokens carry more meaning?
  2. 2Should agents discover context on demand, or should systems anticipate what they will need?
  3. 3What deserves to become durable memory rather than remaining a trace, artifact, or temporary task state?
  4. 4Can an agent remain coherent across sessions without turning summaries into an unreliable substitute for state?
  5. 5How can we evaluate whether an agent remembers usefully rather than merely remembering more?

The Sessions

1Session 1start here

The Context Budget Is the Architecture

The problem

A larger context window appears to promise a simple solution: keep more history, tool output, and reference material available to the model. But the readings challenge that intuition from different directions. Anthropic defines context engineering as the continual curation of the tokens available at inference time, while Google argues that context should be treated as a compiled, per-call view over richer underlying state.

Google’s earlier account of Gemini’s long-context work supplies an important counterpoint. Larger windows genuinely enable new behaviors, including reasoning over large codebases, long videos, and specialized reference material. Yet the later engineering perspective suggests that raw capacity does not eliminate relevance, cost, latency, or “lost in the middle” problems.

OpenAI’s Codex experience makes the issue concrete: repository knowledge became more useful when organized as a navigable map rather than a giant instruction manual. The tension is not between short and long context in the abstract, but between context as accumulated history and context as deliberately structured evidence. That distinction leads directly to the next question: who decides what enters the window?

2Session 2Retrieval or Anticipation?3 readings3Session 3What Should Survive?3 readings4Session 4Crossing the Window3 readings5Session 5Trusting the Residue3 readings