문서 파싱과 레이아웃 이해
PDF와 스캔본을 구조화된 데이터로 되돌린다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Takes a document — a PDF page image or a scan — and outputs structured content: paragraphs, heading levels, tables, figure captions, and the correct reading order. Unlike OCR it does more than transcribe: it decides what each block is and which comes first. Unlike information extraction it first rebuilds the layout skeleton rather than targeting specific fields.
기술적으로 구현하는 방법
A typical pipeline segments the page into regions — text blocks, headings, tables, figures — with detection or segmentation, then handles each: text blocks go through recognition, tables recover row-column and merged-cell structure, and formulas and charts are modelled separately. A layout model predicts block order, and multi-column or cross-column content needs dedicated sorting. Another route hands the whole page to a multimodal model that emits structured markup directly.
대표 제품
7GPT-4o
2024네이티브 멀티모달 범용 모델. 텍스트·이미지·오디오를 한 창구에서 다룬다
Gemini
2023네이티브 멀티모달에 초장문 문맥을 다루는 범용 모델
Claude
2023긴 문맥과 안전 정렬로 알려진 범용 대화 모델
Qwen
2023여러 규모와 멀티모달 버전을 아우르는 오픈웨이트 모델 계열
ERNIE
2019지식 강화 사전학습에서 출발한 중국어 모델, 초기 대표 버전
Hunyuan
2023오픈웨이트 버전을 포함한 텐센트의 범용 모델 계열
NotebookLM
2023제공한 자료만 근거로 답하고 출처를 하나씩 제시한다
관련 기관
대표적 용도
- Structuring invoices, contracts and forms
- Table extraction from filings and research reports
- Digitising archives and case files
- Paper ingestion and knowledge-base building
성능을 평가하는 방법
- Layout element F1
- Correctness of region classification such as heading, table and body
- Table TEDS
- Similarity of table structure trees, measuring row-column recovery
- Reading-order accuracy
- Agreement of the block sequence with the document logic
경계와 난점
- Cross-page tables and merged cells are hard to recover, misaligning rows or losing hierarchy
- Handwritten notes, stamps and stickers disrupt segmentation and cause missed content
- Skew, perspective and binding shadows in scans make column and table boundaries misjudged