Извлечение информации и распознавание сущностей
Извлечь имена, места и связи из свободного текста
Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ
Takes natural-language text and returns structured fragments: typed entities (person, organisation, date, amount) plus the relations or events among them. Unlike classification it does not give one label to the whole text but locates spans; unlike schema-driven structured output, the extracted targets come from the text itself rather than a fully prescribed field list.
Как это устроено
Early systems used conditional random fields or rule-based sequence labelling, tagging tokens BIO-style; pre-trained encoders with a tagging head then became standard. Relation and event extraction are often framed as entity-pair classification or as generation of triples directly. More recently, large models perform few-shot or zero-shot extraction against a given schema, avoiding per-type annotation.
Примеры продуктов
5GPT-4o
2024Универсальная модель с нативной мультимодальностью: текст, изображение и звук через один вход
Qwen
2023Семейство с открытыми весами, охватывающее разные размеры и мультимодальные версии
ERNIE
2019Китайская модель, начавшая с обогащённого знаниями предобучения, ранняя заметная версия
Doubao
2023Универсальная диалоговая модель и приложение ByteDance
GLM
2023Китайская универсальная модель, начавшая с авторегрессивного заполнения пропусков
Связанные организации
Типичное применение
- Clause and amount extraction from contracts and filings
- Résumé parsing and talent-pool building
- Drug and symptom recognition in clinical notes
- News events and knowledge-graph construction
Как её оценивают
- Span-level F1
- A hit requires both boundary and type to be correct
- Relation F1
- Share of triples (head, relation, tail) matched exactly
- Exact-match rate
- Share of records whose fields are all correct
Границы и трудности
- Nested and overlapping entities are flattened by token-level tagging schemes
- Cross-sentence coreference is hard; pronouns bind to the wrong antecedent
- Domain terms and novel words are missed when the type was unseen in training
Концепции в основе
Токенизация
Модель читает не символы, а токены — способ разбиения тихо определяет и возможности, и стоимость
Предобучение и дообучение
Сначала выучить язык на огромных неразмеченных корпусах, затем специализироваться на малых данных — самый экономный по данным подход в современном ИИ
Обучение с учителем
Пары «вопрос — ответ» учат модель отвечать самостоятельно