Optical Character Recognition
Read the text in an image as editable characters
WHAT THIS CAPABILITY MEANS
Takes an image containing text — a scan, a photo, a screenshot — and outputs the character sequence, usually with location boxes. It reads characters rather than interpreting them: the output is a transcription, not a meaning. Unlike document parsing, which also cares about layout such as tables, columns and reading order, OCR is only responsible for getting the characters right.
How it is done
The classic pipeline has two steps: a detection network finds quadrilateral boxes for lines or words, then each crop is passed to a sequence recogniser that decodes characters. Recognition moved from CNN plus recurrent layers with connectionist temporal classification to attention decoders and plain convolutional or Transformer designs. More recently, multimodal models transcribe end to end and handle irregular layouts better.
Representative products
5GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Gemini
2023A natively multimodal general model built for very long context
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
ERNIE
2019A Chinese model that began with knowledge-enhanced pretraining, an early landmark version
Hunyuan
2023Tencent’s general model family, with open-weight versions
Organizations involved
Typical uses
- Field capture from invoices, receipts and IDs
- Digitising paper archives and books
- Street-sign and licence-plate reading
- Copying text from screenshots and photos
How it is evaluated
- Character error rate
- Substitutions, deletions and insertions over the truth length
- Word error rate
- Word-level error rate, sensitive to segmentation
- Detection F1
- Localisation accuracy of text regions
Limits and hard parts
- Handwriting and poor scans — blurry, skewed, smudged — drive the error rate up sharply
- Vertical text, mixed scripts and decorative fonts are frequently missed or misread
- Tables and formulas come out as a character stream, losing row-column structure and super/subscripts
Concepts behind it
Image Representation
To a machine, a photo is nothing but stacked grids of numbers
Convolutional Neural Networks
Replacing full connections with “look locally, reuse the same filter everywhere” — the idea that made image recognition work
Tokenization
Models do not read characters, they read tokens — and how you split text quietly sets both capability and cost