मुख्य सामग्री पर जाएँ

दस्तावेज़ पार्सिंग और लेआउट समझ

PDF और स्कैन को संरचित डेटा में बदलना

डेटा और दस्तावेज़मध्यवर्ती #38
इनपुटइमेजटेबल

यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।

यह क्षमता क्या है

Takes a document — a PDF page image or a scan — and outputs structured content: paragraphs, heading levels, tables, figure captions, and the correct reading order. Unlike OCR it does more than transcribe: it decides what each block is and which comes first. Unlike information extraction it first rebuilds the layout skeleton rather than targeting specific fields.

तकनीकी रूप से कैसे

A typical pipeline segments the page into regions — text blocks, headings, tables, figures — with detection or segmentation, then handles each: text blocks go through recognition, tables recover row-column and merged-cell structure, and formulas and charts are modelled separately. A layout model predicts block order, and multi-column or cross-column content needs dedicated sorting. Another route hands the whole page to a multimodal model that emits structured markup directly.

प्रतिनिधि उत्पाद

7

GPT-4o

2024
OpenAI

मूल रूप से बहुविध सामान्य मॉडल — पाठ, चित्र और ऑडियो एक ही द्वार से

मॉडल बंद स्रोत
टेक्स्टइमेजऑडियोटेक्स्टऑडियो

Gemini

2023
Google DeepMind

मूल रूप से बहुविध, अति-लंबे संदर्भ के लिए बना सामान्य मॉडल

मॉडल बंद स्रोत
टेक्स्टइमेजऑडियोवीडियोटेक्स्ट

Claude

2023
Anthropic

लंबे संदर्भ और सुरक्षा-संरेखण के लिए जाना जाने वाला सामान्य संवाद मॉडल

मॉडल बंद स्रोत
टेक्स्टइमेजटेक्स्ट

Qwen

2023
Alibaba (Qwen)

अनेक आकारों और बहुविध संस्करणों वाला ओपन-वेट परिवार

मॉडल खुले वेट
टेक्स्टइमेजटेक्स्ट

ERNIE

2019
Baidu

ज्ञान-संवर्धित प्रीट्रेनिंग से शुरू हुआ चीनी मॉडल, शुरुआती प्रतिनिधि संस्करण

मॉडल बंद स्रोत
टेक्स्टटेक्स्ट

Hunyuan

2023
Tencent (Hunyuan)

ओपन-वेट संस्करणों सहित टेनसेंट का सामान्य मॉडल परिवार

मॉडल खुले वेट
टेक्स्टइमेजटेक्स्ट

NotebookLM

2023
Google DeepMind

केवल दिए गए स्रोतों से उत्तर देता है, हवाले के साथ

ऐप फ़्रीमियम
टेक्स्टटेबलऑडियोटेक्स्टऑडियो

संबंधित संस्थान

सामान्य उपयोग

  • Structuring invoices, contracts and forms
  • Table extraction from filings and research reports
  • Digitising archives and case files
  • Paper ingestion and knowledge-base building

इसका मूल्यांकन कैसे होता है

Layout element F1
Correctness of region classification such as heading, table and body
Table TEDS
Similarity of table structure trees, measuring row-column recovery
Reading-order accuracy
Agreement of the block sequence with the document logic

सीमाएँ और कठिनाइयाँ

  • Cross-page tables and merged cells are hard to recover, misaligning rows or losing hierarchy
  • Handwritten notes, stamps and stickers disrupt segmentation and cause missed content
  • Skew, perspective and binding shadows in scans make column and table boundaries misjudged

इसके पीछे की अवधारणाएँ