Разбор документов и вёрстки
Превратить PDF и сканы в структурированные данные
Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ
Takes a document — a PDF page image or a scan — and outputs structured content: paragraphs, heading levels, tables, figure captions, and the correct reading order. Unlike OCR it does more than transcribe: it decides what each block is and which comes first. Unlike information extraction it first rebuilds the layout skeleton rather than targeting specific fields.
Как это устроено
A typical pipeline segments the page into regions — text blocks, headings, tables, figures — with detection or segmentation, then handles each: text blocks go through recognition, tables recover row-column and merged-cell structure, and formulas and charts are modelled separately. A layout model predicts block order, and multi-column or cross-column content needs dedicated sorting. Another route hands the whole page to a multimodal model that emits structured markup directly.
Примеры продуктов
7GPT-4o
2024Универсальная модель с нативной мультимодальностью: текст, изображение и звук через один вход
Gemini
2023Универсальная модель с нативной мультимодальностью, рассчитанная на очень длинный контекст
Claude
2023Универсальная диалоговая модель, известная длинным контекстом и выравниванием по безопасности
Qwen
2023Семейство с открытыми весами, охватывающее разные размеры и мультимодальные версии
ERNIE
2019Китайская модель, начавшая с обогащённого знаниями предобучения, ранняя заметная версия
Hunyuan
2023Семейство универсальных моделей Tencent с открытыми весами
NotebookLM
2023Отвечает только по вашим источникам и ссылается на них
Связанные организации
Типичное применение
- Structuring invoices, contracts and forms
- Table extraction from filings and research reports
- Digitising archives and case files
- Paper ingestion and knowledge-base building
Как её оценивают
- Layout element F1
- Correctness of region classification such as heading, table and body
- Table TEDS
- Similarity of table structure trees, measuring row-column recovery
- Reading-order accuracy
- Agreement of the block sequence with the document logic
Границы и трудности
- Cross-page tables and merged cells are hard to recover, misaligning rows or losing hierarchy
- Handwritten notes, stamps and stickers disrupt segmentation and cause missed content
- Skew, perspective and binding shadows in scans make column and table boundaries misjudged
Концепции в основе
Цифровое представление изображения
Для машины фотография — лишь набор наложенных сеток чисел
Механизм внимания
Каждая позиция может напрямую «видеть» все остальные и динамически распределять внимание по релевантности
Архитектура Transformer
Замена эстафеты по одному слову залом, где все говорят сразу, — и дальние зависимости оказываются в одном шаге