Phân tích tài liệu và bố cục
Chuyển PDF và bản quét thành dữ liệu có cấu trúc
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
NĂNG LỰC NÀY NGHĨA LÀ GÌ
Takes a document — a PDF page image or a scan — and outputs structured content: paragraphs, heading levels, tables, figure captions, and the correct reading order. Unlike OCR it does more than transcribe: it decides what each block is and which comes first. Unlike information extraction it first rebuilds the layout skeleton rather than targeting specific fields.
Làm ra sao về mặt kỹ thuật
A typical pipeline segments the page into regions — text blocks, headings, tables, figures — with detection or segmentation, then handles each: text blocks go through recognition, tables recover row-column and merged-cell structure, and formulas and charts are modelled separately. A layout model predicts block order, and multi-column or cross-column content needs dedicated sorting. Another route hands the whole page to a multimodal model that emits structured markup directly.
Sản phẩm tiêu biểu
7GPT-4o
2024Mô hình đa phương thức gốc, xử lý văn bản, hình ảnh và âm thanh qua một cửa vào
Gemini
2023Mô hình đa phương thức gốc, xử lý ngữ cảnh siêu dài
Claude
2023Mô hình hội thoại phổ thông nổi bật với ngữ cảnh dài và căn chỉnh an toàn
Qwen
2023Dòng trọng số mở với nhiều kích cỡ và bản đa phương thức
ERNIE
2019Mô hình tiếng Trung khởi đầu bằng tiền huấn luyện tăng cường tri thức, bản đầu tiên tiêu biểu
Hunyuan
2023Dòng mô hình phổ thông của Tencent, có bản trọng số mở
NotebookLM
2023Chỉ trả lời từ nguồn bạn đưa, kèm trích dẫn từng chỗ
Tổ chức liên quan
Cách dùng tiêu biểu
- Structuring invoices, contracts and forms
- Table extraction from filings and research reports
- Digitising archives and case files
- Paper ingestion and knowledge-base building
Đánh giá nó tốt hay không thế nào
- Layout element F1
- Correctness of region classification such as heading, table and body
- Table TEDS
- Similarity of table structure trees, measuring row-column recovery
- Reading-order accuracy
- Agreement of the block sequence with the document logic
Ranh giới và điểm khó
- Cross-page tables and merged cells are hard to recover, misaligning rows or losing hierarchy
- Handwritten notes, stamps and stickers disrupt segmentation and cause missed content
- Skew, perspective and binding shadows in scans make column and table boundaries misjudged
Các khái niệm đằng sau
Biểu diễn số của ảnh
Với máy móc, một bức ảnh chỉ là các lưới số xếp chồng
Cơ chế chú ý (Attention)
Mọi vị trí đều có thể nhìn thẳng vào mọi vị trí khác và phân bổ chú ý theo mức liên quan
Kiến trúc Transformer
Thay cách truyền từng từ bằng một phòng họp nơi mọi từ cùng lên tiếng, để phụ thuộc xa chỉ còn cách một bước