The Signal Behind “I Don’t Know”
The problem
The first difficulty is conceptual: what would count as a system recognizing a knowledge gap? A model can produce a confidence score, refuse to answer, or assign a probability to its own correctness without possessing a robust, domain-general notion of uncertainty. The useful distinction is between surface confidence, behavioral calibration, and information encoded in internal states.
Anthropic’s work on self-evaluation provides an early positive result: sufficiently large models can estimate whether their answers are correct and can sometimes predict whether they know an answer before producing one. OpenAI’s work on verbalized uncertainty approaches the same problem differently, showing that calibration can be expressed in natural language and can survive some distribution shift. Google DeepMind’s “Experts Don’t Cheat” pushes further by separating ordinary response variability from epistemic uncertainty—the gap between a model’s learned approximation and the underlying process.
Together, these readings establish both the promise and the fragility of self-knowledge. The signal may exist before the model is asked to report it, but calibration depends heavily on task format, training objectives, and the distinction between uncertainty that reflects the world and uncertainty that reflects ignorance. That leads to the next question: if useful warning signals exist internally, can we actually locate and interpret them?