テキスト分類と感情分析
テキストに既定のラベルを付ける
本ページの本文は英語で提供されています。タイトルと導入は日本語化されています。
この能力とは何か
Takes a piece of text and outputs its category — sentiment polarity, topic, intent or policy violation. The label set is fixed at training time, so it answers which class rather than what is said. A single text can carry several labels, and the model can also emit a probability distribution over the label set.
技術的にどう実現するか
Early systems paired bag-of-words with linear classifiers; fine-tuning a pre-trained encoder such as the BERT family later became the standard when labelled data is scarce. A classification head sits on top of the encoder and is trained with cross-entropy. With many labels and few examples, it can be recast as sentence-pair entailment or done zero-shot by prompting a large model.
代表的な製品
5GPT-4o
2024ネイティブにマルチモーダルな汎用モデル。テキスト・画像・音声をひとつの入口で扱う
Qwen
2023多様な規模とマルチモーダル版を備えたオープンウェイトのモデル群
Hunyuan
2023開放ウェイト版を含むテンセントの汎用モデル群
ERNIE
2019知識増強の事前学習から始まった中国語モデル、その初期の代表的版
Gemini
2024検索・オフィス・マルチモーダルモデルを一つの対話入口に集約
関連する組織
代表的な用途
- Sentiment polarity of reviews and public opinion
- Spam and policy-violation filtering
- Routing tickets by topic and priority
- Intent detection to drive dialogue
どう評価するか
- F1
- Harmonic mean of precision and recall; more trustworthy than accuracy under imbalance
- ROC-AUC
- Ranking quality, threshold-independent
- Confusion matrix
- Per-class error direction, to locate systematic bias
限界と難しさ
- Under heavy class imbalance, minority classes are ignored while accuracy still looks high
- Wording drift outside the training domain degrades predictions
- Sarcasm, negation and double negation are often flipped, since the literal words point the other way