Ology

From Next-Token Prediction to Trustworthy Machines

Large language models: how they learn and generate text, why they seem intelligent, uses, limitations, reliability, and safety questions

5 sessions · 15 readings · 7 views

Curated by Rajeev Chhajer

The Predictor That Learned Too Muchsession 1Teaching a Predictor to Pleasesession 2Intelligence, Error, and the Hidden Processsession 3From Chatbot to Operating Systemsession 4Measuring What We Cannot Reliably Seesession 5

Opening

Large language models are now best understood not as isolated chatbots, but as adaptable components inside systems that write code, retrieve information, use tools, reason over long tasks, and sometimes act with limited autonomy. The important shift is not simply that models have become larger. Training objectives, post-training methods, inference-time reasoning, open-weight releases, and application architecture have all changed what these systems can do.

That progress makes the central picture harder, not easier. A model trained to predict text can produce explanations, programs, arguments, and apparently deliberate plans. Yet the same system may confidently invent a citation, exploit an evaluation loophole, or fail when a task requires grounding in the world rather than fluency. Recent reasoning models add another layer: more computation at inference time can improve difficult problem-solving, while creating new questions about transparency, cost, and control.

The readings therefore move from capability to behavior, then from behavior to systems and governance. They compare scaling with data, alignment with imitation, reasoning with hallucination, and autonomy with reliability. Their authors often disagree—not only about what works, but about what should count as evidence that it works.

The seminar’s central problem is this: how can we understand and build systems whose impressive linguistic competence remains useful, inspectable, and dependable when the world demands more than plausible text?

Five questions worth arguing about

  1. 1Does next-token prediction explain LLM competence, or merely describe a useful interface to deeper learned structure?
  2. 2Do post-training methods align models with human intent, or mostly teach them to perform helpfulness convincingly?
  3. 3When models reason longer, are they becoming more reliable thinkers or more effective generators of plausible mistakes?
  4. 4Should practical LLM systems be designed around autonomous agents, or around simpler supervised workflows with tools?
  5. 5Can evaluation and monitoring make increasingly capable models governable, or do they move faster than our measurements?

The Sessions

1Session 1start here

The Predictor That Learned Too Much

The problem

The beginner’s explanation of an LLM is deceptively simple: it predicts the next token. The difficulty is explaining how that objective yields broad factual knowledge, translation, coding, summarization, and coherent interaction without treating “prediction” as a magic word.

Google DeepMind’s Gopher report is useful precisely because it separates capabilities that improve with scale from those that remain stubbornly weak. Its discussion of dialogue, fact-checking, reasoning, repetition, stereotypes, and confident errors presents a model that is neither empty autocomplete nor a general-purpose mind. The Chinchilla work then complicates the idea that “bigger” is the main route to better models: data, parameters, and compute must be balanced, and a smaller model trained on more data can outperform a much larger one.

The GPT-4 research post provides a bridge from laboratory scaling to general-purpose use. It documents broad competence while retaining an explicit account of unreliability. Together, these readings ask whether language modeling gives rise to intelligence, simulates many of its outward forms, or makes that distinction less useful than we expect. The next session follows the model after pretraining, when developers try to turn capability into a cooperative assistant.

2Session 2Teaching a Predictor to Please3 readings3Session 3Intelligence, Error, and the Hidden Process3 readings4Session 4From Chatbot to Operating System3 readings5Session 5Measuring What We Cannot Reliably See3 readings
From Next-Token Prediction to Trustworthy Machines · Ology