The Document Is Not the Text
The problem
The first mistake in building a knowledge system is to assume that extraction is mainly transcription. PDFs encode visual layout rather than semantic structure; spreadsheets encode meaning through headers, formulas, and neighboring cells; web pages mix navigation, prose, tables, and metadata. A system that extracts every character correctly can still produce a semantically corrupted representation.
Hugging Face’s Document AI survey frames parsing as a collection of distinct problems—OCR, layout analysis, table detection, table structure recognition, and document question answering. OpenAI’s contract workflow shows the operational consequence: scanned contracts and handwritten edits must become structured fields while retaining annotations and references for human review. The two readings therefore pull in different directions: one emphasizes model and task decomposition, while the other emphasizes a reviewable business workflow.
Anthropic’s Contextual Retrieval adds a second complication. Even after extraction, chunking can erase the document-level context needed to interpret a passage. Its proposed remedy is not full symbolic normalization, but contextual enrichment before hybrid retrieval. The session establishes the first design tension: should a knowledge pipeline preserve the original artifact, reconstruct a normalized intermediate representation, or maintain both?