Optische Zeichenerkennung
Text im Bild als bearbeitbare Zeichen lesen
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes an image containing text — a scan, a photo, a screenshot — and outputs the character sequence, usually with location boxes. It reads characters rather than interpreting them: the output is a transcription, not a meaning. Unlike document parsing, which also cares about layout such as tables, columns and reading order, OCR is only responsible for getting the characters right.
Wie sie technisch umgesetzt wird
The classic pipeline has two steps: a detection network finds quadrilateral boxes for lines or words, then each crop is passed to a sequence recogniser that decodes characters. Recognition moved from CNN plus recurrent layers with connectionist temporal classification to attention decoders and plain convolutional or Transformer designs. More recently, multimodal models transcribe end to end and handle irregular layouts better.
Repräsentative Produkte
5GPT-4o
2024Ein von Grund auf multimodales Allzweckmodell: Text, Bild und Audio über einen Zugang
Gemini
2023Ein von Grund auf multimodales Allzweckmodell für sehr langen Kontext
Qwen
2023Eine Open-Weights-Familie über viele Größen hinweg, mit multimodalen Versionen
ERNIE
2019Ein chinesisches Modell, das mit wissensverstärktem Vortraining begann – frühe prägende Version
Hunyuan
2023Tencent Generalmodell-Familie mit Open-Weights-Versionen
Beteiligte Organisationen
Typische Verwendungen
- Field capture from invoices, receipts and IDs
- Digitising paper archives and books
- Street-sign and licence-plate reading
- Copying text from screenshots and photos
Wie sie bewertet wird
- Character error rate
- Substitutions, deletions and insertions over the truth length
- Word error rate
- Word-level error rate, sensitive to segmentation
- Detection F1
- Localisation accuracy of text regions
Grenzen und schwierige Punkte
- Handwriting and poor scans — blurry, skewed, smudged — drive the error rate up sharply
- Vertical text, mixed scripts and decorative fonts are frequently missed or misread
- Tables and formulas come out as a character stream, losing row-column structure and super/subscripts
Konzepte dahinter
Digitale Bilddarstellung
Für eine Maschine ist ein Foto nichts als gestapelte Zahlenraster
Faltungsnetze
Vollverbindungen durch „lokal schauen, überall denselben Filter wiederverwenden“ ersetzen – die Idee, die Bilderkennung erst brauchbar machte
Tokenisierung
Modelle lesen keine Zeichen, sondern Tokens – und die Zerlegung bestimmt still Fähigkeit wie Kosten