The Predictor That Learned Too Much
The problem
The beginner’s explanation of an LLM is deceptively simple: it predicts the next token. The difficulty is explaining how that objective yields broad factual knowledge, translation, coding, summarization, and coherent interaction without treating “prediction” as a magic word.
Google DeepMind’s Gopher report is useful precisely because it separates capabilities that improve with scale from those that remain stubbornly weak. Its discussion of dialogue, fact-checking, reasoning, repetition, stereotypes, and confident errors presents a model that is neither empty autocomplete nor a general-purpose mind. The Chinchilla work then complicates the idea that “bigger” is the main route to better models: data, parameters, and compute must be balanced, and a smaller model trained on more data can outperform a much larger one.
The GPT-4 research post provides a bridge from laboratory scaling to general-purpose use. It documents broad competence while retaining an explicit account of unreliability. Together, these readings ask whether language modeling gives rise to intelligence, simulates many of its outward forms, or makes that distinction less useful than we expect. The next session follows the model after pretraining, when developers try to turn capability into a cooperative assistant.
Readings