Nhận dạng ký tự quang học
Đọc chữ trong ảnh thành ký tự có thể chỉnh sửa
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
NĂNG LỰC NÀY NGHĨA LÀ GÌ
Takes an image containing text — a scan, a photo, a screenshot — and outputs the character sequence, usually with location boxes. It reads characters rather than interpreting them: the output is a transcription, not a meaning. Unlike document parsing, which also cares about layout such as tables, columns and reading order, OCR is only responsible for getting the characters right.
Làm ra sao về mặt kỹ thuật
The classic pipeline has two steps: a detection network finds quadrilateral boxes for lines or words, then each crop is passed to a sequence recogniser that decodes characters. Recognition moved from CNN plus recurrent layers with connectionist temporal classification to attention decoders and plain convolutional or Transformer designs. More recently, multimodal models transcribe end to end and handle irregular layouts better.
Sản phẩm tiêu biểu
5GPT-4o
2024Mô hình đa phương thức gốc, xử lý văn bản, hình ảnh và âm thanh qua một cửa vào
Gemini
2023Mô hình đa phương thức gốc, xử lý ngữ cảnh siêu dài
Qwen
2023Dòng trọng số mở với nhiều kích cỡ và bản đa phương thức
ERNIE
2019Mô hình tiếng Trung khởi đầu bằng tiền huấn luyện tăng cường tri thức, bản đầu tiên tiêu biểu
Hunyuan
2023Dòng mô hình phổ thông của Tencent, có bản trọng số mở
Tổ chức liên quan
Cách dùng tiêu biểu
- Field capture from invoices, receipts and IDs
- Digitising paper archives and books
- Street-sign and licence-plate reading
- Copying text from screenshots and photos
Đánh giá nó tốt hay không thế nào
- Character error rate
- Substitutions, deletions and insertions over the truth length
- Word error rate
- Word-level error rate, sensitive to segmentation
- Detection F1
- Localisation accuracy of text regions
Ranh giới và điểm khó
- Handwriting and poor scans — blurry, skewed, smudged — drive the error rate up sharply
- Vertical text, mixed scripts and decorative fonts are frequently missed or misread
- Tables and formulas come out as a character stream, losing row-column structure and super/subscripts
Các khái niệm đằng sau
Biểu diễn số của ảnh
Với máy móc, một bức ảnh chỉ là các lưới số xếp chồng
Mạng nơ-ron tích chập
Thay kết nối đầy đủ bằng “nhìn cục bộ, dùng lại cùng một thước ở mọi nơi” — ý tưởng đưa nhận dạng ảnh thành khả thi
Token hóa
Mô hình không đọc ký tự mà đọc token; cách tách từ âm thầm quyết định năng lực và chi phí