Trích xuất thông tin và nhận dạng thực thể
Trích tên, địa điểm và quan hệ từ văn bản tự do
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
NĂNG LỰC NÀY NGHĨA LÀ GÌ
Takes natural-language text and returns structured fragments: typed entities (person, organisation, date, amount) plus the relations or events among them. Unlike classification it does not give one label to the whole text but locates spans; unlike schema-driven structured output, the extracted targets come from the text itself rather than a fully prescribed field list.
Làm ra sao về mặt kỹ thuật
Early systems used conditional random fields or rule-based sequence labelling, tagging tokens BIO-style; pre-trained encoders with a tagging head then became standard. Relation and event extraction are often framed as entity-pair classification or as generation of triples directly. More recently, large models perform few-shot or zero-shot extraction against a given schema, avoiding per-type annotation.
Sản phẩm tiêu biểu
5GPT-4o
2024Mô hình đa phương thức gốc, xử lý văn bản, hình ảnh và âm thanh qua một cửa vào
Qwen
2023Dòng trọng số mở với nhiều kích cỡ và bản đa phương thức
ERNIE
2019Mô hình tiếng Trung khởi đầu bằng tiền huấn luyện tăng cường tri thức, bản đầu tiên tiêu biểu
Doubao
2023Mô hình hội thoại phổ thông và ứng dụng của ByteDance
GLM
2023Mô hình phổ thông tiếng Trung khởi đầu bằng tiền huấn luyện điền khuyết tự hồi quy
Tổ chức liên quan
Cách dùng tiêu biểu
- Clause and amount extraction from contracts and filings
- Résumé parsing and talent-pool building
- Drug and symptom recognition in clinical notes
- News events and knowledge-graph construction
Đánh giá nó tốt hay không thế nào
- Span-level F1
- A hit requires both boundary and type to be correct
- Relation F1
- Share of triples (head, relation, tail) matched exactly
- Exact-match rate
- Share of records whose fields are all correct
Ranh giới và điểm khó
- Nested and overlapping entities are flattened by token-level tagging schemes
- Cross-sentence coreference is hard; pronouns bind to the wrong antecedent
- Domain terms and novel words are missed when the type was unseen in training
Các khái niệm đằng sau
Token hóa
Mô hình không đọc ký tự mà đọc token; cách tách từ âm thầm quyết định năng lực và chi phí
Tiền huấn luyện và tinh chỉnh
Học ngôn ngữ trước từ lượng lớn văn bản không nhãn, rồi chuyên biệt hóa với ít dữ liệu — mô hình tiết kiệm dữ liệu nhất của AI hiện đại
Học có giám sát
Từng cặp câu hỏi – đáp án dạy mô hình tự trả lời