Ology

The Agent That Knows What It Knows

AI agents track and update claims by provenance, confidence, uncertainty, and temporal validity, improving answer trustworthiness

5 sessions · 16 readings · 1 view

Curated by Ben Lorenzo

The Citation Is Not the Claimsession 1From Context Window to Knowledge Systemsession 2Memory Must Be Maintained, Not Merely Retrievedsession 3Can We Test Knowledge Drift?session 4Trust Without Overtrustsession 5

Opening

AI agents are moving from answering isolated questions to maintaining working context across documents, tools, databases, and weeks of interaction. That shift changes the reliability problem. It is no longer enough for an answer to sound plausible or even to cite a source; the system must know which claim came from where, when it was true, how strongly it is supported, and whether later evidence has superseded it.

Recent systems make parts of this possible. Anthropic has turned source attribution into an API capability, OpenAI describes agents that combine metadata, lineage, institutional knowledge, runtime inspection, and editable memory, and Google DeepMind has explored provenance as a way to help people judge context rather than merely detect generated content. (anthropic.com)

But provenance is not verification, confidence is not calibration, and memory is not a database with embeddings. A citation may identify a document without proving that it supports the precise claim. A remembered rule may have been correct only for one schema version, one user, or one date. A system that updates aggressively may amplify mistakes; one that updates cautiously may remain stale.

This seminar follows the problem from source-traceable claims to maintainable knowledge, then to evaluation, control, and human trust. Its central question is whether an agent can become more useful by exposing the limits and history of what it knows—or whether visible provenance merely makes an unreliable system look more defensible.

Five questions worth arguing about

  1. 1When does a citation make an agent’s claim defensible rather than merely traceable?
  2. 2Should durable agent knowledge be organized as retrieved context, structured memory, or an auditable system of record?
  3. 3Can agents update knowledge safely without turning uncertainty and contradiction into persistent memory?
  4. 4Which evaluations reveal whether an agent’s knowledge remains correct as sources, users, and environments change?
  5. 5Does greater transparency actually improve reliance, or can provenance create unjustified confidence in agent-generated answers?

The Sessions

1Session 1start here

The Citation Is Not the Claim

The problem

The first temptation is to treat provenance as a presentation feature: attach a source to each sentence and call the answer trustworthy. Anthropic’s Citations API makes that approach concrete by allowing models to cite passages from supplied documents, while its web-search tooling extends the idea to current external information. The engineering achievement is real, but it leaves open a harder question: does source linkage establish support, or only make inspection possible? (anthropic.com)

Google DeepMind’s work on Backstory complicates the picture from another direction. It distinguishes whether an image was generated from whether it is trustworthy, emphasizing origin, reuse, alteration, metadata, and surrounding context. The older Verifiable Data Audit proposal adds a systems-level conception of trust: records should be complete, tamper-evident, and inspectable, not merely displayed after the fact. (deepmind.google)

Together, these readings separate at least three ideas that are often collapsed: provenance, evidential support, and auditability. The next session asks what kind of knowledge infrastructure is needed if claims are to remain inspectable after the initial answer has been generated.

2Session 2From Context Window to Knowledge System3 readings3Session 3Memory Must Be Maintained, Not Merely Retrieved3 readings4Session 4Can We Test Knowledge Drift?3 readings5Session 5Trust Without Overtrust3 readings