Textklassifikation und Stimmungsanalyse
Einem Text ein vordefiniertes Label zuweisen
Der vollständige Artikel liegt auf Englisch vor; Titel und Zusammenfassung sind lokalisiert.
WAS DIESE FÄHIGKEIT BEDEUTET
Takes a piece of text and outputs its category — sentiment polarity, topic, intent or policy violation. The label set is fixed at training time, so it answers which class rather than what is said. A single text can carry several labels, and the model can also emit a probability distribution over the label set.
Wie sie technisch umgesetzt wird
Early systems paired bag-of-words with linear classifiers; fine-tuning a pre-trained encoder such as the BERT family later became the standard when labelled data is scarce. A classification head sits on top of the encoder and is trained with cross-entropy. With many labels and few examples, it can be recast as sentence-pair entailment or done zero-shot by prompting a large model.
Repräsentative Produkte
5GPT-4o
2024Ein von Grund auf multimodales Allzweckmodell: Text, Bild und Audio über einen Zugang
Qwen
2023Eine Open-Weights-Familie über viele Größen hinweg, mit multimodalen Versionen
Hunyuan
2023Tencent Generalmodell-Familie mit Open-Weights-Versionen
ERNIE
2019Ein chinesisches Modell, das mit wissensverstärktem Vortraining begann – frühe prägende Version
Gemini
2024Ein Chat-Eingang, der Suche, Office-Apps und ein multimodales Modell bündelt
Beteiligte Organisationen
Typische Verwendungen
- Sentiment polarity of reviews and public opinion
- Spam and policy-violation filtering
- Routing tickets by topic and priority
- Intent detection to drive dialogue
Wie sie bewertet wird
- F1
- Harmonic mean of precision and recall; more trustworthy than accuracy under imbalance
- ROC-AUC
- Ranking quality, threshold-independent
- Confusion matrix
- Per-class error direction, to locate systematic bias
Grenzen und schwierige Punkte
- Under heavy class imbalance, minority classes are ignored while accuracy still looks high
- Wording drift outside the training domain degrades predictions
- Sarcasm, negation and double negation are often flipped, since the literal words point the other way
Konzepte dahinter
Überwachtes Lernen
Paare aus Frage und Antwort lehren das Modell, selbst zu antworten
Wort-Embeddings
Wörter werden zu Koordinaten: Synonyme rücken zusammen und Bedeutung lässt sich erstmals addieren
Modellbewertung und Kreuzvalidierung
Genauigkeit ist die täuschendste Metrik – falsch ausgewertet, bricht alles Weitere zusammen