Reconhecimento óptico de caracteres
Ler o texto de uma imagem como caracteres editáveis
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE ESTA CAPACIDADE SIGNIFICA
Takes an image containing text — a scan, a photo, a screenshot — and outputs the character sequence, usually with location boxes. It reads characters rather than interpreting them: the output is a transcription, not a meaning. Unlike document parsing, which also cares about layout such as tables, columns and reading order, OCR is only responsible for getting the characters right.
Como é feita tecnicamente
The classic pipeline has two steps: a detection network finds quadrilateral boxes for lines or words, then each crop is passed to a sequence recogniser that decodes characters. Recognition moved from CNN plus recurrent layers with connectionist temporal classification to attention decoders and plain convolutional or Transformer designs. More recently, multimodal models transcribe end to end and handle irregular layouts better.
Produtos representativos
5GPT-4o
2024Um modelo geral nativamente multimodal: texto, imagem e áudio por uma única porta
Gemini
2023Um modelo geral nativamente multimodal, feito para contextos muito longos
Qwen
2023Uma família de pesos abertos com muitos tamanhos e versões multimodais
ERNIE
2019Um modelo chinês que começou com pré-treinamento enriquecido por conhecimento, versão inicial representativa
Hunyuan
2023A família de modelos gerais da Tencent, com versões de pesos abertos
Organizações relacionadas
Usos típicos
- Field capture from invoices, receipts and IDs
- Digitising paper archives and books
- Street-sign and licence-plate reading
- Copying text from screenshots and photos
Como avaliar se funciona bem
- Character error rate
- Substitutions, deletions and insertions over the truth length
- Word error rate
- Word-level error rate, sensitive to segmentation
- Detection F1
- Localisation accuracy of text regions
Limites e dificuldades
- Handwriting and poor scans — blurry, skewed, smudged — drive the error rate up sharply
- Vertical text, mixed scripts and decorative fonts are frequently missed or misread
- Tables and formulas come out as a character stream, losing row-column structure and super/subscripts
Conceitos por trás
Representação digital de imagens
Para uma máquina, uma foto não passa de grades de números sobrepostas
Redes neurais convolucionais
Substituir as conexões densas por “olhar localmente e reutilizar o mesmo filtro em toda parte”: a ideia que tornou o reconhecimento de imagens funcional
Tokenização
Os modelos não leem caracteres, leem tokens; e como você divide define silenciosamente a capacidade e o custo