Text Classification & Sentiment
Assign a predefined label to a piece of text
WHAT THIS CAPABILITY MEANS
Takes a piece of text and outputs its category — sentiment polarity, topic, intent or policy violation. The label set is fixed at training time, so it answers which class rather than what is said. A single text can carry several labels, and the model can also emit a probability distribution over the label set.
How it is done
Early systems paired bag-of-words with linear classifiers; fine-tuning a pre-trained encoder such as the BERT family later became the standard when labelled data is scarce. A classification head sits on top of the encoder and is trained with cross-entropy. With many labels and few examples, it can be recast as sentence-pair entailment or done zero-shot by prompting a large model.
Representative products
5GPT-4o
2024A natively multimodal general model, with text, image and audio through one door
Qwen
2023An open-weight family spanning many sizes, with multimodal versions
Hunyuan
2023Tencent’s general model family, with open-weight versions
ERNIE
2019A Chinese model that began with knowledge-enhanced pretraining, an early landmark version
Gemini
2024One chat entry point that gathers search, office apps and a multimodal model
Organizations involved
Typical uses
- Sentiment polarity of reviews and public opinion
- Spam and policy-violation filtering
- Routing tickets by topic and priority
- Intent detection to drive dialogue
How it is evaluated
- F1
- Harmonic mean of precision and recall; more trustworthy than accuracy under imbalance
- ROC-AUC
- Ranking quality, threshold-independent
- Confusion matrix
- Per-class error direction, to locate systematic bias
Limits and hard parts
- Under heavy class imbalance, minority classes are ignored while accuracy still looks high
- Wording drift outside the training domain degrades predictions
- Sarcasm, negation and double negation are often flipped, since the literal words point the other way
Concepts behind it
Supervised Learning
Pairs of questions and answers teach a model to answer on its own
Word Embeddings
Turning words into coordinates — synonyms land near each other, and meaning becomes something you can add and subtract
Model Evaluation & Cross-Validation
Accuracy is the easiest metric to fool you — get evaluation wrong and everything else follows