Ology

Before the Failure: How AI Systems Learn to Notice Their Own Limits

AI self-monitoring methods for detecting knowledge gaps and emerging failures before they cause harm

5 sessions · 16 readings · 1 view

Curated by Ben Lorenzo

The Signal Behind “I Don’t Know”session 1Reading the Silent Statesession 2Watching the Reasoning Before the Actsession 3When Uncertainty Has to Change the Plansession 4The Generalization Barriersession 5

Opening

As AI systems become more autonomous, reliability increasingly depends on more than producing good outputs on average. A system must recognize when its internal evidence is weak, when its reasoning is going off track, and when a familiar-looking situation is actually outside the conditions under which it learned to behave safely.

Recent work makes this possibility less speculative. Language models can sometimes estimate whether their answers are correct, expose latent concepts that never appear in their outputs, and reveal internal signals associated with hallucination, deception, or unsafe reasoning. Autonomous-driving systems provide a complementary perspective: uncertainty is valuable only when it changes what the system does, such as slowing down, deferring, or entering a safe fallback mode.

But self-monitoring is not the same as self-understanding. A readable representation may be merely correlated with a behavior; verbalized confidence may fail under distribution shift; a monitor may catch current failures while missing future ones; and a safety case may establish operational evidence without revealing the model’s internal causes. The central challenge is therefore not simply whether AI systems contain useful warning signals, but whether those signals can be made causal, calibrated, actionable, and transferable.

The five sessions investigate that challenge from the inside out: first asking what it means for a model to know its limits, then examining its representations, monitoring its reasoning, connecting those ideas to autonomous driving, and finally testing whether self-monitoring can generalize.

Five questions worth arguing about

  1. 1Can a model’s confidence signal reveal genuine knowledge gaps, or merely confidence in a familiar response format?
  2. 2When hidden representations are readable, do they explain behavior—or only provide a persuasive after-the-fact correlate?
  3. 3Does monitoring reasoning expose emerging failure risks, or encourage systems to learn how to hide them?
  4. 4Can autonomous driving’s uncertainty and safety-case practices detect danger early enough to justify action?
  5. 5What must remain stable for self-monitoring methods to generalize from one model, task, or domain to another?

The Sessions

1Session 1start here

The Signal Behind “I Don’t Know”

The problem

The first difficulty is conceptual: what would count as a system recognizing a knowledge gap? A model can produce a confidence score, refuse to answer, or assign a probability to its own correctness without possessing a robust, domain-general notion of uncertainty. The useful distinction is between surface confidence, behavioral calibration, and information encoded in internal states.

Anthropic’s work on self-evaluation provides an early positive result: sufficiently large models can estimate whether their answers are correct and can sometimes predict whether they know an answer before producing one. OpenAI’s work on verbalized uncertainty approaches the same problem differently, showing that calibration can be expressed in natural language and can survive some distribution shift. Google DeepMind’s “Experts Don’t Cheat” pushes further by separating ordinary response variability from epistemic uncertainty—the gap between a model’s learned approximation and the underlying process.

Together, these readings establish both the promise and the fragility of self-knowledge. The signal may exist before the model is asked to report it, but calibration depends heavily on task format, training objectives, and the distinction between uncertainty that reflects the world and uncertainty that reflects ignorance. That leads to the next question: if useful warning signals exist internally, can we actually locate and interpret them?

2Session 2Reading the Silent State4 readings3Session 3Watching the Reasoning Before the Act3 readings4Session 4When Uncertainty Has to Change the Plan3 readings5Session 5The Generalization Barrier3 readings