التعرّف الضوئي على الحروف
قراءة النص في الصورة كحروف قابلة للتحرير
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
ما الذي تعنيه هذه القدرة
Takes an image containing text — a scan, a photo, a screenshot — and outputs the character sequence, usually with location boxes. It reads characters rather than interpreting them: the output is a transcription, not a meaning. Unlike document parsing, which also cares about layout such as tables, columns and reading order, OCR is only responsible for getting the characters right.
كيف تُنفَّذ تقنيًا
The classic pipeline has two steps: a detection network finds quadrilateral boxes for lines or words, then each crop is passed to a sequence recogniser that decodes characters. Recognition moved from CNN plus recurrent layers with connectionist temporal classification to attention decoders and plain convolutional or Transformer designs. More recently, multimodal models transcribe end to end and handle irregular layouts better.
منتجات تمثيلية
5GPT-4o
2024نموذج عام متعدد الوسائط بطبيعته، يجمع النص والصورة والصوت في مدخل واحد
Gemini
2023نموذج عام متعدد الوسائط بطبيعته، مصمم لسياقات طويلة جدًا
Qwen
2023عائلة بأوزان مفتوحة تغطي أحجامًا متعددة، مع نسخ متعددة الوسائط
ERNIE
2019نموذج صيني بدأ بالتدريب المسبق المعزّز بالمعرفة، نسخة مبكرة ممثلة
Hunyuan
2023عائلة نماذج تينسنت العامة، مع نسخ بأوزان مفتوحة
المؤسسات ذات الصلة
الاستخدامات الشائعة
- Field capture from invoices, receipts and IDs
- Digitising paper archives and books
- Street-sign and licence-plate reading
- Copying text from screenshots and photos
كيف يُقاس مدى جودتها
- Character error rate
- Substitutions, deletions and insertions over the truth length
- Word error rate
- Word-level error rate, sensitive to segmentation
- Detection F1
- Localisation accuracy of text regions
الحدود والصعوبات
- Handwriting and poor scans — blurry, skewed, smudged — drive the error rate up sharply
- Vertical text, mixed scripts and decorative fonts are frequently missed or misread
- Tables and formulas come out as a character stream, losing row-column structure and super/subscripts
المفاهيم الكامنة وراءها
التمثيل الرقمي للصورة
الصورة بالنسبة للآلة ليست سوى شبكات من الأرقام متراكبة
الشبكات العصبية الالتفافية
استبدال الاتصال الكامل بـ«النظر محليًا وإعادة استخدام المسطرة نفسها في كل مكان» — الفكرة التي جعلت التعرّف على الصور ممكنًا
الترميز إلى رموز (Tokenization)
لا تقرأ النماذج الحروف بل الرموز، وطريقة التقسيم تحدّد القدرة والكلفة بصمت