Análise de documentos e layout
Converter PDFs e digitalizações em dados estruturados
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE ESTA CAPACIDADE SIGNIFICA
Takes a document — a PDF page image or a scan — and outputs structured content: paragraphs, heading levels, tables, figure captions, and the correct reading order. Unlike OCR it does more than transcribe: it decides what each block is and which comes first. Unlike information extraction it first rebuilds the layout skeleton rather than targeting specific fields.
Como é feita tecnicamente
A typical pipeline segments the page into regions — text blocks, headings, tables, figures — with detection or segmentation, then handles each: text blocks go through recognition, tables recover row-column and merged-cell structure, and formulas and charts are modelled separately. A layout model predicts block order, and multi-column or cross-column content needs dedicated sorting. Another route hands the whole page to a multimodal model that emits structured markup directly.
Produtos representativos
7GPT-4o
2024Um modelo geral nativamente multimodal: texto, imagem e áudio por uma única porta
Gemini
2023Um modelo geral nativamente multimodal, feito para contextos muito longos
Claude
2023Um modelo de conversa geral conhecido por contexto longo e alinhamento de segurança
Qwen
2023Uma família de pesos abertos com muitos tamanhos e versões multimodais
ERNIE
2019Um modelo chinês que começou com pré-treinamento enriquecido por conhecimento, versão inicial representativa
Hunyuan
2023A família de modelos gerais da Tencent, com versões de pesos abertos
NotebookLM
2023Responde apenas com as fontes que você fornece, com citações
Organizações relacionadas
Usos típicos
- Structuring invoices, contracts and forms
- Table extraction from filings and research reports
- Digitising archives and case files
- Paper ingestion and knowledge-base building
Como avaliar se funciona bem
- Layout element F1
- Correctness of region classification such as heading, table and body
- Table TEDS
- Similarity of table structure trees, measuring row-column recovery
- Reading-order accuracy
- Agreement of the block sequence with the document logic
Limites e dificuldades
- Cross-page tables and merged cells are hard to recover, misaligning rows or losing hierarchy
- Handwritten notes, stamps and stickers disrupt segmentation and cause missed content
- Skew, perspective and binding shadows in scans make column and table boundaries misjudged
Conceitos por trás
Representação digital de imagens
Para uma máquina, uma foto não passa de grades de números sobrepostas
Mecanismo de atenção
Cada posição pode olhar diretamente para todas as outras e distribuir atenção conforme a relevância
Arquitetura Transformer
Substituir o revezamento palavra a palavra por uma sala onde todos falam ao mesmo tempo, deixando dependências distantes a um salto