Reconocimiento óptico de caracteres
Leer el texto de una imagen como caracteres editables
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ SIGNIFICA ESTA CAPACIDAD
Takes an image containing text — a scan, a photo, a screenshot — and outputs the character sequence, usually with location boxes. It reads characters rather than interpreting them: the output is a transcription, not a meaning. Unlike document parsing, which also cares about layout such as tables, columns and reading order, OCR is only responsible for getting the characters right.
Cómo se consigue técnicamente
The classic pipeline has two steps: a detection network finds quadrilateral boxes for lines or words, then each crop is passed to a sequence recogniser that decodes characters. Recognition moved from CNN plus recurrent layers with connectionist temporal classification to attention decoders and plain convolutional or Transformer designs. More recently, multimodal models transcribe end to end and handle irregular layouts better.
Productos representativos
5GPT-4o
2024Un modelo general nativamente multimodal: texto, imagen y audio por una misma puerta
Gemini
2023Un modelo general nativamente multimodal, pensado para contextos muy largos
Qwen
2023Una familia de pesos abiertos con muchos tamaños y versiones multimodales
ERNIE
2019Un modelo chino que comenzó con preentrenamiento enriquecido con conocimiento, versión temprana representativa
Hunyuan
2023La familia de modelos generales de Tencent, con versiones de pesos abiertos
Organizaciones relacionadas
Usos típicos
- Field capture from invoices, receipts and IDs
- Digitising paper archives and books
- Street-sign and licence-plate reading
- Copying text from screenshots and photos
Cómo se evalúa
- Character error rate
- Substitutions, deletions and insertions over the truth length
- Word error rate
- Word-level error rate, sensitive to segmentation
- Detection F1
- Localisation accuracy of text regions
Límites y dificultades
- Handwriting and poor scans — blurry, skewed, smudged — drive the error rate up sharply
- Vertical text, mixed scripts and decorative fonts are frequently missed or misread
- Tables and formulas come out as a character stream, losing row-column structure and super/subscripts
Conceptos detrás
Representación digital de imágenes
Para una máquina, una foto no es más que cuadrículas de números superpuestas
Redes neuronales convolucionales
Sustituir las conexiones completas por “mirar en local y reutilizar la misma regla en todas partes”: la idea que hizo funcionar el reconocimiento de imágenes
Tokenización
Los modelos no leen caracteres, leen tokens; y cómo los dividas decide en silencio capacidad y coste