Ology

From Messy Files to Living Knowledge

AI-driven systems transform messy sources into interconnected knowledge bases, tracking provenance, uncertainty, and knowledge gaps.

5 sessions · 20 readings · 1 view

Curated by Ben Lorenzo

The Document Is Not the Textsession 1Names, Schemas, and False Matchessession 2Graphs Without the Graph Hypesession 3Every Answer Has a Historysession 4The Knowledge Base Must Know Its Gapssession 5

Opening

AI knowledge systems are moving beyond “search over documents.” The harder ambition is to turn PDFs, spreadsheets, databases, code, and web pages into a shared substrate that can be queried, updated, audited, and trusted. Recent systems increasingly combine document parsing, schema metadata, retrieval, structured extraction, graph construction, tool use, and evaluation rather than treating knowledge as a pile of text chunks.

That shift matters because the most damaging failures often occur before generation. A PDF parser can destroy table structure; a schema mapper can confuse similar fields; an entity resolver can merge two people or split one organization; a retrieval system can surface a stale policy while presenting it as current. Better language models help, but they do not eliminate these data and systems problems.

The emerging designs also disagree about what the “knowledge base” should be. Some favor contextualized text and hybrid retrieval; others build explicit graphs, ontologies, or executable metadata layers. Some preserve uncertainty and conflicting claims, while others normalize toward a single operational representation. The central question is not whether AI can extract knowledge, but how much structure, evidence, and revision machinery a system needs before its answers deserve trust.

The five sessions follow that transformation from ingestion to maintenance: preserving structure, resolving meaning, choosing representations, tracing claims through transformations, and evaluating whether the resulting system knows what it does not know.

Five questions worth arguing about

  1. 1Should document pipelines optimize for faithful preservation or aggressive normalization before language models ever see the data?
  2. 2Can schema and entity resolution be automated without quietly turning ambiguity into false certainty?
  3. 3When does a knowledge graph improve reasoning enough to justify its construction and maintenance costs?
  4. 4Is provenance useful only after an answer fails, or should it shape every intermediate representation?
  5. 5How should a living knowledge system represent stale, conflicting, missing, or weakly supported beliefs?

The Sessions

1Session 1start here

The Document Is Not the Text

The problem

The first mistake in building a knowledge system is to assume that extraction is mainly transcription. PDFs encode visual layout rather than semantic structure; spreadsheets encode meaning through headers, formulas, and neighboring cells; web pages mix navigation, prose, tables, and metadata. A system that extracts every character correctly can still produce a semantically corrupted representation.

Hugging Face’s Document AI survey frames parsing as a collection of distinct problems—OCR, layout analysis, table detection, table structure recognition, and document question answering. OpenAI’s contract workflow shows the operational consequence: scanned contracts and handwritten edits must become structured fields while retaining annotations and references for human review. The two readings therefore pull in different directions: one emphasizes model and task decomposition, while the other emphasizes a reviewable business workflow.

Anthropic’s Contextual Retrieval adds a second complication. Even after extraction, chunking can erase the document-level context needed to interpret a passage. Its proposed remedy is not full symbolic normalization, but contextual enrichment before hybrid retrieval. The session establishes the first design tension: should a knowledge pipeline preserve the original artifact, reconstruct a normalized intermediate representation, or maintain both?

2Session 2Names, Schemas, and False Matches4 readings3Session 3Graphs Without the Graph Hype4 readings4Session 4Every Answer Has a History4 readings5Session 5The Knowledge Base Must Know Its Gaps4 readings
From Messy Files to Living Knowledge · Ology