Informationsextraktion und NER
Namen, Orte und Relationen aus freiem Text ziehen
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes natural-language text and returns structured fragments: typed entities (person, organisation, date, amount) plus the relations or events among them. Unlike classification it does not give one label to the whole text but locates spans; unlike schema-driven structured output, the extracted targets come from the text itself rather than a fully prescribed field list.
Wie sie technisch umgesetzt wird
Early systems used conditional random fields or rule-based sequence labelling, tagging tokens BIO-style; pre-trained encoders with a tagging head then became standard. Relation and event extraction are often framed as entity-pair classification or as generation of triples directly. More recently, large models perform few-shot or zero-shot extraction against a given schema, avoiding per-type annotation.
Repräsentative Produkte
5GPT-4o
2024Ein von Grund auf multimodales Allzweckmodell: Text, Bild und Audio über einen Zugang
Qwen
2023Eine Open-Weights-Familie über viele Größen hinweg, mit multimodalen Versionen
ERNIE
2019Ein chinesisches Modell, das mit wissensverstärktem Vortraining begann – frühe prägende Version
Doubao
2023ByteDances allgemeines Dialogmodell und App
GLM
2023Ein chinesisches Allzweckmodell, das mit autoregressivem Lückenfüllen begann
Beteiligte Organisationen
Typische Verwendungen
- Clause and amount extraction from contracts and filings
- Résumé parsing and talent-pool building
- Drug and symptom recognition in clinical notes
- News events and knowledge-graph construction
Wie sie bewertet wird
- Span-level F1
- A hit requires both boundary and type to be correct
- Relation F1
- Share of triples (head, relation, tail) matched exactly
- Exact-match rate
- Share of records whose fields are all correct
Grenzen und schwierige Punkte
- Nested and overlapping entities are flattened by token-level tagging schemes
- Cross-sentence coreference is hard; pronouns bind to the wrong antecedent
- Domain terms and novel words are missed when the type was unseen in training
Konzepte dahinter
Tokenisierung
Modelle lesen keine Zeichen, sondern Tokens – und die Zerlegung bestimmt still Fähigkeit wie Kosten
Vor- und Feintraining
Erst Sprache aus riesigem ungelabelten Text lernen, dann mit wenig Daten spezialisieren – das dateneffizienteste Paradigma der modernen KI
Überwachtes Lernen
Paare aus Frage und Antwort lehren das Modell, selbst zu antworten