Análisis de documentos y maquetación
Convertir PDF y escaneos en datos estructurados
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ SIGNIFICA ESTA CAPACIDAD
Takes a document — a PDF page image or a scan — and outputs structured content: paragraphs, heading levels, tables, figure captions, and the correct reading order. Unlike OCR it does more than transcribe: it decides what each block is and which comes first. Unlike information extraction it first rebuilds the layout skeleton rather than targeting specific fields.
Cómo se consigue técnicamente
A typical pipeline segments the page into regions — text blocks, headings, tables, figures — with detection or segmentation, then handles each: text blocks go through recognition, tables recover row-column and merged-cell structure, and formulas and charts are modelled separately. A layout model predicts block order, and multi-column or cross-column content needs dedicated sorting. Another route hands the whole page to a multimodal model that emits structured markup directly.
Productos representativos
7GPT-4o
2024Un modelo general nativamente multimodal: texto, imagen y audio por una misma puerta
Gemini
2023Un modelo general nativamente multimodal, pensado para contextos muy largos
Claude
2023Un modelo de chat general conocido por su contexto largo y su alineación de seguridad
Qwen
2023Una familia de pesos abiertos con muchos tamaños y versiones multimodales
ERNIE
2019Un modelo chino que comenzó con preentrenamiento enriquecido con conocimiento, versión temprana representativa
Hunyuan
2023La familia de modelos generales de Tencent, con versiones de pesos abiertos
NotebookLM
2023Responde solo con las fuentes que le das, citando cada una
Organizaciones relacionadas
Usos típicos
- Structuring invoices, contracts and forms
- Table extraction from filings and research reports
- Digitising archives and case files
- Paper ingestion and knowledge-base building
Cómo se evalúa
- Layout element F1
- Correctness of region classification such as heading, table and body
- Table TEDS
- Similarity of table structure trees, measuring row-column recovery
- Reading-order accuracy
- Agreement of the block sequence with the document logic
Límites y dificultades
- Cross-page tables and merged cells are hard to recover, misaligning rows or losing hierarchy
- Handwritten notes, stamps and stickers disrupt segmentation and cause missed content
- Skew, perspective and binding shadows in scans make column and table boundaries misjudged
Conceptos detrás
Representación digital de imágenes
Para una máquina, una foto no es más que cuadrículas de números superpuestas
Mecanismo de atención
Cada posición puede mirar directamente a todas las demás y repartir su atención según la relevancia
Arquitectura Transformer
Sustituir el relé palabra a palabra por una sala donde todos hablan a la vez, para que las dependencias lejanas estén a un salto