Classificação de texto e sentimento
Atribuir um rótulo predefinido a um texto
O texto completo é apresentado em inglês; o título e o resumo estão traduzidos.
O QUE ESTA CAPACIDADE SIGNIFICA
Takes a piece of text and outputs its category — sentiment polarity, topic, intent or policy violation. The label set is fixed at training time, so it answers which class rather than what is said. A single text can carry several labels, and the model can also emit a probability distribution over the label set.
Como é feita tecnicamente
Early systems paired bag-of-words with linear classifiers; fine-tuning a pre-trained encoder such as the BERT family later became the standard when labelled data is scarce. A classification head sits on top of the encoder and is trained with cross-entropy. With many labels and few examples, it can be recast as sentence-pair entailment or done zero-shot by prompting a large model.
Produtos representativos
5GPT-4o
2024Um modelo geral nativamente multimodal: texto, imagem e áudio por uma única porta
Qwen
2023Uma família de pesos abertos com muitos tamanhos e versões multimodais
Hunyuan
2023A família de modelos gerais da Tencent, com versões de pesos abertos
ERNIE
2019Um modelo chinês que começou com pré-treinamento enriquecido por conhecimento, versão inicial representativa
Gemini
2024Uma entrada de chat que reúne busca, escritório e um modelo multimodal
Organizações relacionadas
Usos típicos
- Sentiment polarity of reviews and public opinion
- Spam and policy-violation filtering
- Routing tickets by topic and priority
- Intent detection to drive dialogue
Como avaliar se funciona bem
- F1
- Harmonic mean of precision and recall; more trustworthy than accuracy under imbalance
- ROC-AUC
- Ranking quality, threshold-independent
- Confusion matrix
- Per-class error direction, to locate systematic bias
Limites e dificuldades
- Under heavy class imbalance, minority classes are ignored while accuracy still looks high
- Wording drift outside the training domain degrades predictions
- Sarcasm, negation and double negation are often flipped, since the literal words point the other way
Conceitos por trás
Aprendizagem supervisionada
Pares de pergunta e resposta ensinam o modelo a responder sozinho
Embeddings de palavras
Transformar palavras em coordenadas: sinônimos se agrupam sozinhos e o significado passa a poder ser somado e subtraído
Avaliação de modelos e validação cruzada
A acurácia é a métrica mais fácil de enganar — erre na avaliação e tudo o mais desmorona