Document Parsing & Layout
Turn PDFs and scans into structured data
WHAT THIS CAPABILITY MEANS
Takes a document — a PDF page image or a scan — and outputs structured content: paragraphs, heading levels, tables, figure captions, and the correct reading order. Unlike OCR it does more than transcribe: it decides what each block is and which comes first. Unlike information extraction it first rebuilds the layout skeleton rather than targeting specific fields.
How it is done
A typical pipeline segments the page into regions — text blocks, headings, tables, figures — with detection or segmentation, then handles each: text blocks go through recognition, tables recover row-column and merged-cell structure, and formulas and charts are modelled separately. A layout model predicts block order, and multi-column or cross-column content needs dedicated sorting. Another route hands the whole page to a multimodal model that emits structured markup directly.
Representative products
7GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Gemini
2023A natively multimodal general model built for very long context
Claude
2023A general chat model known for long context and safety alignment
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
ERNIE
2019A Chinese model that began with knowledge-enhanced pretraining, an early landmark version
Hunyuan
2023Tencent’s general model family, with open-weight versions
NotebookLM
2023Answers only from the sources you give it, with citations
Organizations involved
Typical uses
- Structuring invoices, contracts and forms
- Table extraction from filings and research reports
- Digitising archives and case files
- Paper ingestion and knowledge-base building
How it is evaluated
- Layout element F1
- Correctness of region classification such as heading, table and body
- Table TEDS
- Similarity of table structure trees, measuring row-column recovery
- Reading-order accuracy
- Agreement of the block sequence with the document logic
Limits and hard parts
- Cross-page tables and merged cells are hard to recover, misaligning rows or losing hierarchy
- Handwritten notes, stamps and stickers disrupt segmentation and cause missed content
- Skew, perspective and binding shadows in scans make column and table boundaries misjudged
Concepts behind it
Image Representation
To a machine, a photo is nothing but stacked grids of numbers
Attention Mechanism
Every position can look directly at every other position and dynamically weight how much attention to pay
Transformer Architecture
Replacing word-by-word relay with a room where everyone speaks at once, so long-range dependencies are one hop away