Перейти к содержимому
Атлас ИИ

Разбор документов и вёрстки

Превратить PDF и сканы в структурированные данные

Данные и документыСредний #38
входИзображениеТаблица

Полный текст статьи представлен на английском; заголовок и аннотация локализованы.

ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ

Takes a document — a PDF page image or a scan — and outputs structured content: paragraphs, heading levels, tables, figure captions, and the correct reading order. Unlike OCR it does more than transcribe: it decides what each block is and which comes first. Unlike information extraction it first rebuilds the layout skeleton rather than targeting specific fields.

Как это устроено

A typical pipeline segments the page into regions — text blocks, headings, tables, figures — with detection or segmentation, then handles each: text blocks go through recognition, tables recover row-column and merged-cell structure, and formulas and charts are modelled separately. A layout model predicts block order, and multi-column or cross-column content needs dedicated sorting. Another route hands the whole page to a multimodal model that emits structured markup directly.

Примеры продуктов

7

GPT-4o

2024
OpenAI

Универсальная модель с нативной мультимодальностью: текст, изображение и звук через один вход

Модель Закрытый
ТекстИзображениеАудиоТекстАудио

Gemini

2023
Google DeepMind

Универсальная модель с нативной мультимодальностью, рассчитанная на очень длинный контекст

Модель Закрытый
ТекстИзображениеАудиоВидеоТекст

Claude

2023
Anthropic

Универсальная диалоговая модель, известная длинным контекстом и выравниванием по безопасности

Модель Закрытый
ТекстИзображениеТекст

Qwen

2023
Alibaba (Qwen)

Семейство с открытыми весами, охватывающее разные размеры и мультимодальные версии

Модель Открытые веса
ТекстИзображениеТекст

ERNIE

2019
Baidu

Китайская модель, начавшая с обогащённого знаниями предобучения, ранняя заметная версия

Модель Закрытый
ТекстТекст

Hunyuan

2023
Tencent (Hunyuan)

Семейство универсальных моделей Tencent с открытыми весами

Модель Открытые веса
ТекстИзображениеТекст

NotebookLM

2023
Google DeepMind

Отвечает только по вашим источникам и ссылается на них

Приложение Freemium
ТекстТаблицаАудиоТекстАудио

Связанные организации

Типичное применение

  • Structuring invoices, contracts and forms
  • Table extraction from filings and research reports
  • Digitising archives and case files
  • Paper ingestion and knowledge-base building

Как её оценивают

Layout element F1
Correctness of region classification such as heading, table and body
Table TEDS
Similarity of table structure trees, measuring row-column recovery
Reading-order accuracy
Agreement of the block sequence with the document logic

Границы и трудности

  • Cross-page tables and merged cells are hard to recover, misaligning rows or losing hierarchy
  • Handwritten notes, stamps and stickers disrupt segmentation and cause missed content
  • Skew, perspective and binding shadows in scans make column and table boundaries misjudged

Концепции в основе