تحليل المستندات وفهم التخطيط
تحويل ملفات PDF والممسوحات إلى بيانات منظّمة
يُعرض النص الكامل باللغة الإنجليزية؛ وقد تمت ترجمة العنوان والملخص.
ما الذي تعنيه هذه القدرة
Takes a document — a PDF page image or a scan — and outputs structured content: paragraphs, heading levels, tables, figure captions, and the correct reading order. Unlike OCR it does more than transcribe: it decides what each block is and which comes first. Unlike information extraction it first rebuilds the layout skeleton rather than targeting specific fields.
كيف تُنفَّذ تقنيًا
A typical pipeline segments the page into regions — text blocks, headings, tables, figures — with detection or segmentation, then handles each: text blocks go through recognition, tables recover row-column and merged-cell structure, and formulas and charts are modelled separately. A layout model predicts block order, and multi-column or cross-column content needs dedicated sorting. Another route hands the whole page to a multimodal model that emits structured markup directly.
منتجات تمثيلية
7GPT-4o
2024نموذج عام متعدد الوسائط بطبيعته، يجمع النص والصورة والصوت في مدخل واحد
Gemini
2023نموذج عام متعدد الوسائط بطبيعته، مصمم لسياقات طويلة جدًا
Claude
2023نموذج محادثة عام يشتهر بسياقه الطويل ومواءمته الأمنية
Qwen
2023عائلة بأوزان مفتوحة تغطي أحجامًا متعددة، مع نسخ متعددة الوسائط
ERNIE
2019نموذج صيني بدأ بالتدريب المسبق المعزّز بالمعرفة، نسخة مبكرة ممثلة
Hunyuan
2023عائلة نماذج تينسنت العامة، مع نسخ بأوزان مفتوحة
NotebookLM
2023يجيب من المصادر التي تقدّمها فقط، مع ذكر المراجع
المؤسسات ذات الصلة
الاستخدامات الشائعة
- Structuring invoices, contracts and forms
- Table extraction from filings and research reports
- Digitising archives and case files
- Paper ingestion and knowledge-base building
كيف يُقاس مدى جودتها
- Layout element F1
- Correctness of region classification such as heading, table and body
- Table TEDS
- Similarity of table structure trees, measuring row-column recovery
- Reading-order accuracy
- Agreement of the block sequence with the document logic
الحدود والصعوبات
- Cross-page tables and merged cells are hard to recover, misaligning rows or losing hierarchy
- Handwritten notes, stamps and stickers disrupt segmentation and cause missed content
- Skew, perspective and binding shadows in scans make column and table boundaries misjudged
المفاهيم الكامنة وراءها
التمثيل الرقمي للصورة
الصورة بالنسبة للآلة ليست سوى شبكات من الأرقام متراكبة
آلية الانتباه
يستطيع كل موضع أن ينظر مباشرة إلى جميع المواضع الأخرى ويوزّع الانتباه حسب الصلة
معمارية Transformer
يستبدل النقل كلمةً بكلمة بغرفة يتحدث فيها الجميع معاً، فتصبح التبعيات البعيدة على مسافة خطوة واحدة