मुख्य सामग्री पर जाएँ

प्रकाशिक अक्षर पहचान

छवि के पाठ को संपादन-योग्य अक्षरों में पढ़ना

दृष्टि बोधप्रारंभिक #15
इनपुटइमेजटेक्स्ट

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

यह क्षमता क्या है

Takes an image containing text — a scan, a photo, a screenshot — and outputs the character sequence, usually with location boxes. It reads characters rather than interpreting them: the output is a transcription, not a meaning. Unlike document parsing, which also cares about layout such as tables, columns and reading order, OCR is only responsible for getting the characters right.

तकनीकी रूप से कैसे

The classic pipeline has two steps: a detection network finds quadrilateral boxes for lines or words, then each crop is passed to a sequence recogniser that decodes characters. Recognition moved from CNN plus recurrent layers with connectionist temporal classification to attention decoders and plain convolutional or Transformer designs. More recently, multimodal models transcribe end to end and handle irregular layouts better.

प्रतिनिधि उत्पाद

5

संबंधित संस्थान

सामान्य उपयोग

  • Field capture from invoices, receipts and IDs
  • Digitising paper archives and books
  • Street-sign and licence-plate reading
  • Copying text from screenshots and photos

इसका मूल्यांकन कैसे होता है

Character error rate
Substitutions, deletions and insertions over the truth length
Word error rate
Word-level error rate, sensitive to segmentation
Detection F1
Localisation accuracy of text regions

सीमाएँ और कठिनाइयाँ

  • Handwriting and poor scans — blurry, skewed, smudged — drive the error rate up sharply
  • Vertical text, mixed scripts and decorative fonts are frequently missed or misread
  • Tables and formulas come out as a character stream, losing row-column structure and super/subscripts

इसके पीछे की अवधारणाएँ