Классификация текста и анализ тональности
Присвоить тексту заранее заданную метку
Полный текст статьи представлен на английском; заголовок и аннотация локализованы.
ЧТО ЭТО ЗА ВОЗМОЖНОСТЬ
Takes a piece of text and outputs its category — sentiment polarity, topic, intent or policy violation. The label set is fixed at training time, so it answers which class rather than what is said. A single text can carry several labels, and the model can also emit a probability distribution over the label set.
Как это устроено
Early systems paired bag-of-words with linear classifiers; fine-tuning a pre-trained encoder such as the BERT family later became the standard when labelled data is scarce. A classification head sits on top of the encoder and is trained with cross-entropy. With many labels and few examples, it can be recast as sentence-pair entailment or done zero-shot by prompting a large model.
Примеры продуктов
5GPT-4o
2024Универсальная модель с нативной мультимодальностью: текст, изображение и звук через один вход
Qwen
2023Семейство с открытыми весами, охватывающее разные размеры и мультимодальные версии
Hunyuan
2023Семейство универсальных моделей Tencent с открытыми весами
ERNIE
2019Китайская модель, начавшая с обогащённого знаниями предобучения, ранняя заметная версия
Gemini
2024Единое окно диалога, собравшее поиск, офисные приложения и мультимодальную модель
Связанные организации
Типичное применение
- Sentiment polarity of reviews and public opinion
- Spam and policy-violation filtering
- Routing tickets by topic and priority
- Intent detection to drive dialogue
Как её оценивают
- F1
- Harmonic mean of precision and recall; more trustworthy than accuracy under imbalance
- ROC-AUC
- Ranking quality, threshold-independent
- Confusion matrix
- Per-class error direction, to locate systematic bias
Границы и трудности
- Under heavy class imbalance, minority classes are ignored while accuracy still looks high
- Wording drift outside the training domain degrades predictions
- Sarcasm, negation and double negation are often flipped, since the literal words point the other way
Концепции в основе
Обучение с учителем
Пары «вопрос — ответ» учат модель отвечать самостоятельно
Векторные представления слов
Превращение слов в координаты: синонимы сами собираются вместе, и смыслом впервые можно складывать и вычитать
Оценка модели и кросс-валидация
Точность — самый обманчивый показатель: ошибётесь в оценке, и рухнет всё остальное