Classification de texte et analyse de sentiment
Attribuer une étiquette prédéfinie à un texte
Le texte intégral est présenté en anglais ; le titre et le résumé sont localisés.
CE QUE DÉSIGNE CETTE CAPACITÉ
Takes a piece of text and outputs its category — sentiment polarity, topic, intent or policy violation. The label set is fixed at training time, so it answers which class rather than what is said. A single text can carry several labels, and the model can also emit a probability distribution over the label set.
Comment c'est fait
Early systems paired bag-of-words with linear classifiers; fine-tuning a pre-trained encoder such as the BERT family later became the standard when labelled data is scarce. A classification head sits on top of the encoder and is trained with cross-entropy. With many labels and few examples, it can be recast as sentence-pair entailment or done zero-shot by prompting a large model.
Produits représentatifs
5GPT-4o
2024Un modèle général nativement multimodal : texte, image et audio par une même entrée
Qwen
2023Une famille à poids ouverts couvrant de nombreuses tailles, avec des versions multimodales
Hunyuan
2023La famille de modèles généraux de Tencent, avec des versions à poids ouverts
ERNIE
2019Un modèle chinois parti d’un préentraînement enrichi par la connaissance, version phare précoce
Gemini
2024Une entrée de discussion qui réunit recherche, bureautique et modèle multimodal
Organisations concernées
Usages typiques
- Sentiment polarity of reviews and public opinion
- Spam and policy-violation filtering
- Routing tickets by topic and priority
- Intent detection to drive dialogue
Comment on l'évalue
- F1
- Harmonic mean of precision and recall; more trustworthy than accuracy under imbalance
- ROC-AUC
- Ranking quality, threshold-independent
- Confusion matrix
- Per-class error direction, to locate systematic bias
Limites et points difficiles
- Under heavy class imbalance, minority classes are ignored while accuracy still looks high
- Wording drift outside the training domain degrades predictions
- Sarcasm, negation and double negation are often flipped, since the literal words point the other way
Concepts sous-jacents
Apprentissage supervisé
Des paires question-réponse apprennent au modèle à répondre seul
Plongements de mots
Transformer les mots en coordonnées : les synonymes se rapprochent et le sens devient une grandeur que l’on peut additionner
Évaluation de modèle et validation croisée
L’exactitude est la métrique la plus trompeuse — une évaluation erronée fait tout s’effondrer