Extraction d’information et NER
Extraire noms, lieux et relations d’un texte libre
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE DÉSIGNE CETTE CAPACITÉ
Takes natural-language text and returns structured fragments: typed entities (person, organisation, date, amount) plus the relations or events among them. Unlike classification it does not give one label to the whole text but locates spans; unlike schema-driven structured output, the extracted targets come from the text itself rather than a fully prescribed field list.
Comment c'est fait
Early systems used conditional random fields or rule-based sequence labelling, tagging tokens BIO-style; pre-trained encoders with a tagging head then became standard. Relation and event extraction are often framed as entity-pair classification or as generation of triples directly. More recently, large models perform few-shot or zero-shot extraction against a given schema, avoiding per-type annotation.
Produits représentatifs
5GPT-4o
2024Un modèle général nativement multimodal : texte, image et audio par une même entrée
Qwen
2023Une famille à poids ouverts couvrant de nombreuses tailles, avec des versions multimodales
ERNIE
2019Un modèle chinois parti d’un préentraînement enrichi par la connaissance, version phare précoce
Doubao
2023Le modèle de conversation général et l’application de ByteDance
GLM
2023Un modèle général chinois parti d’un préentraînement par remplissage de blancs autorégressif
Organisations concernées
Usages typiques
- Clause and amount extraction from contracts and filings
- Résumé parsing and talent-pool building
- Drug and symptom recognition in clinical notes
- News events and knowledge-graph construction
Comment on l'évalue
- Span-level F1
- A hit requires both boundary and type to be correct
- Relation F1
- Share of triples (head, relation, tail) matched exactly
- Exact-match rate
- Share of records whose fields are all correct
Limites et points difficiles
- Nested and overlapping entities are flattened by token-level tagging schemes
- Cross-sentence coreference is hard; pronouns bind to the wrong antecedent
- Domain terms and novel words are missed when the type was unseen in training
Concepts sous-jacents
Tokenisation
Les modèles ne lisent pas des caractères mais des tokens ; le découpage fixe silencieusement la capacité et le coût
Pré-entraînement et affinage
Apprendre d’abord la langue sur d’immenses corpus non annotés, puis se spécialiser avec peu de données : le paradigme le plus économe en données
Apprentissage supervisé
Des paires question-réponse apprennent au modèle à répondre seul