文書解析とレイアウト理解
PDF やスキャンを構造化データに戻す
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes a document — a PDF page image or a scan — and outputs structured content: paragraphs, heading levels, tables, figure captions, and the correct reading order. Unlike OCR it does more than transcribe: it decides what each block is and which comes first. Unlike information extraction it first rebuilds the layout skeleton rather than targeting specific fields.
技術的にどう実現するか
A typical pipeline segments the page into regions — text blocks, headings, tables, figures — with detection or segmentation, then handles each: text blocks go through recognition, tables recover row-column and merged-cell structure, and formulas and charts are modelled separately. A layout model predicts block order, and multi-column or cross-column content needs dedicated sorting. Another route hands the whole page to a multimodal model that emits structured markup directly.
代表的な製品
7GPT-4o
2024ネイティブにマルチモーダルな汎用モデル。テキスト・画像・音声をひとつの入口で扱う
Gemini
2023ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル
Claude
2023長い文脈と安全性の調整で知られる汎用対話モデル
Qwen
2023多様な規模とマルチモーダル版を備えたオープンウェイトのモデル群
ERNIE
2019知識増強の事前学習から始まった中国語モデル、その初期の代表的版
Hunyuan
2023開放ウェイト版を含むテンセントの汎用モデル群
NotebookLM
2023与えた資料だけを根拠に答え、出典を逐一示す
関連する組織
代表的な用途
- Structuring invoices, contracts and forms
- Table extraction from filings and research reports
- Digitising archives and case files
- Paper ingestion and knowledge-base building
どう評価するか
- Layout element F1
- Correctness of region classification such as heading, table and body
- Table TEDS
- Similarity of table structure trees, measuring row-column recovery
- Reading-order accuracy
- Agreement of the block sequence with the document logic
限界と難しさ
- Cross-page tables and merged cells are hard to recover, misaligning rows or losing hierarchy
- Handwritten notes, stamps and stickers disrupt segmentation and cause missed content
- Skew, perspective and binding shadows in scans make column and table boundaries misjudged