Dokumentanalyse und Layout
PDFs und Scans in strukturierte Daten umwandeln
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes a document — a PDF page image or a scan — and outputs structured content: paragraphs, heading levels, tables, figure captions, and the correct reading order. Unlike OCR it does more than transcribe: it decides what each block is and which comes first. Unlike information extraction it first rebuilds the layout skeleton rather than targeting specific fields.
Wie sie technisch umgesetzt wird
A typical pipeline segments the page into regions — text blocks, headings, tables, figures — with detection or segmentation, then handles each: text blocks go through recognition, tables recover row-column and merged-cell structure, and formulas and charts are modelled separately. A layout model predicts block order, and multi-column or cross-column content needs dedicated sorting. Another route hands the whole page to a multimodal model that emits structured markup directly.
Repräsentative Produkte
7GPT-4o
2024Ein von Grund auf multimodales Allzweckmodell: Text, Bild und Audio über einen Zugang
Gemini
2023Ein von Grund auf multimodales Allzweckmodell für sehr langen Kontext
Claude
2023Ein allgemeines Dialogmodell, bekannt für langen Kontext und Sicherheitsausrichtung
Qwen
2023Eine Open-Weights-Familie über viele Größen hinweg, mit multimodalen Versionen
ERNIE
2019Ein chinesisches Modell, das mit wissensverstärktem Vortraining begann – frühe prägende Version
Hunyuan
2023Tencent Generalmodell-Familie mit Open-Weights-Versionen
NotebookLM
2023Antwortet nur aus den von dir gelieferten Quellen, mit Belegen
Beteiligte Organisationen
Typische Verwendungen
- Structuring invoices, contracts and forms
- Table extraction from filings and research reports
- Digitising archives and case files
- Paper ingestion and knowledge-base building
Wie sie bewertet wird
- Layout element F1
- Correctness of region classification such as heading, table and body
- Table TEDS
- Similarity of table structure trees, measuring row-column recovery
- Reading-order accuracy
- Agreement of the block sequence with the document logic
Grenzen und schwierige Punkte
- Cross-page tables and merged cells are hard to recover, misaligning rows or losing hierarchy
- Handwritten notes, stamps and stickers disrupt segmentation and cause missed content
- Skew, perspective and binding shadows in scans make column and table boundaries misjudged
Konzepte dahinter
Digitale Bilddarstellung
Für eine Maschine ist ein Foto nichts als gestapelte Zahlenraster
Attention-Mechanismus
Jede Position kann direkt auf alle anderen blicken und ihre Aufmerksamkeit nach Relevanz verteilen
Transformer-Architektur
Statt Wort-für-Wort-Stafette ein Raum, in dem alle zugleich sprechen — und weite Abhängigkeiten sind nur einen Schritt entfernt