文字認識(OCR)
画像内の文字を編集可能なテキストにする
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes an image containing text — a scan, a photo, a screenshot — and outputs the character sequence, usually with location boxes. It reads characters rather than interpreting them: the output is a transcription, not a meaning. Unlike document parsing, which also cares about layout such as tables, columns and reading order, OCR is only responsible for getting the characters right.
技術的にどう実現するか
The classic pipeline has two steps: a detection network finds quadrilateral boxes for lines or words, then each crop is passed to a sequence recogniser that decodes characters. Recognition moved from CNN plus recurrent layers with connectionist temporal classification to attention decoders and plain convolutional or Transformer designs. More recently, multimodal models transcribe end to end and handle irregular layouts better.
代表的な製品
5GPT-4o
2024ネイティブにマルチモーダルな汎用モデル。テキスト・画像・音声をひとつの入口で扱う
Gemini
2023ネイティブにマルチモーダルで、超長文脈を扱う汎用モデル
Qwen
2023多様な規模とマルチモーダル版を備えたオープンウェイトのモデル群
ERNIE
2019知識増強の事前学習から始まった中国語モデル、その初期の代表的版
Hunyuan
2023開放ウェイト版を含むテンセントの汎用モデル群
関連する組織
代表的な用途
- Field capture from invoices, receipts and IDs
- Digitising paper archives and books
- Street-sign and licence-plate reading
- Copying text from screenshots and photos
どう評価するか
- Character error rate
- Substitutions, deletions and insertions over the truth length
- Word error rate
- Word-level error rate, sensitive to segmentation
- Detection F1
- Localisation accuracy of text regions
限界と難しさ
- Handwriting and poor scans — blurry, skewed, smudged — drive the error rate up sharply
- Vertical text, mixed scripts and decorative fonts are frequently missed or misread
- Tables and formulas come out as a character stream, losing row-column structure and super/subscripts