Оптическое распознавание символов
Прочитать текст на изображении как редактируемые символы
Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ
Takes an image containing text — a scan, a photo, a screenshot — and outputs the character sequence, usually with location boxes. It reads characters rather than interpreting them: the output is a transcription, not a meaning. Unlike document parsing, which also cares about layout such as tables, columns and reading order, OCR is only responsible for getting the characters right.
Как это устроено
The classic pipeline has two steps: a detection network finds quadrilateral boxes for lines or words, then each crop is passed to a sequence recogniser that decodes characters. Recognition moved from CNN plus recurrent layers with connectionist temporal classification to attention decoders and plain convolutional or Transformer designs. More recently, multimodal models transcribe end to end and handle irregular layouts better.
Примеры продуктов
5GPT-4o
2024Универсальная модель с нативной мультимодальностью: текст, изображение и звук через один вход
Gemini
2023Универсальная модель с нативной мультимодальностью, рассчитанная на очень длинный контекст
Qwen
2023Семейство с открытыми весами, охватывающее разные размеры и мультимодальные версии
ERNIE
2019Китайская модель, начавшая с обогащённого знаниями предобучения, ранняя заметная версия
Hunyuan
2023Семейство универсальных моделей Tencent с открытыми весами
Связанные организации
Типичное применение
- Field capture from invoices, receipts and IDs
- Digitising paper archives and books
- Street-sign and licence-plate reading
- Copying text from screenshots and photos
Как её оценивают
- Character error rate
- Substitutions, deletions and insertions over the truth length
- Word error rate
- Word-level error rate, sensitive to segmentation
- Detection F1
- Localisation accuracy of text regions
Границы и трудности
- Handwriting and poor scans — blurry, skewed, smudged — drive the error rate up sharply
- Vertical text, mixed scripts and decorative fonts are frequently missed or misread
- Tables and formulas come out as a character stream, losing row-column structure and super/subscripts
Концепции в основе
Цифровое представление изображения
Для машины фотография — лишь набор наложенных сеток чисел
Свёрточные нейронные сети
Замена полных связей на «смотреть локально и переиспользовать один и тот же фильтр везде» — идея, сделавшая распознавание изображений рабочим
Токенизация
Модель читает не символы, а токены — способ разбиения тихо определяет и возможности, и стоимость