Clasificación de texto y sentimiento
Asignar una etiqueta predefinida a un texto
El texto completo se presenta en inglés; el título y el resumen están traducidos.
QUÉ SIGNIFICA ESTA CAPACIDAD
Takes a piece of text and outputs its category — sentiment polarity, topic, intent or policy violation. The label set is fixed at training time, so it answers which class rather than what is said. A single text can carry several labels, and the model can also emit a probability distribution over the label set.
Cómo se consigue técnicamente
Early systems paired bag-of-words with linear classifiers; fine-tuning a pre-trained encoder such as the BERT family later became the standard when labelled data is scarce. A classification head sits on top of the encoder and is trained with cross-entropy. With many labels and few examples, it can be recast as sentence-pair entailment or done zero-shot by prompting a large model.
Productos representativos
5GPT-4o
2024Un modelo general nativamente multimodal: texto, imagen y audio por una misma puerta
Qwen
2023Una familia de pesos abiertos con muchos tamaños y versiones multimodales
Hunyuan
2023La familia de modelos generales de Tencent, con versiones de pesos abiertos
ERNIE
2019Un modelo chino que comenzó con preentrenamiento enriquecido con conocimiento, versión temprana representativa
Gemini
2024Una entrada de chat que reúne búsqueda, ofimática y un modelo multimodal
Organizaciones relacionadas
Usos típicos
- Sentiment polarity of reviews and public opinion
- Spam and policy-violation filtering
- Routing tickets by topic and priority
- Intent detection to drive dialogue
Cómo se evalúa
- F1
- Harmonic mean of precision and recall; more trustworthy than accuracy under imbalance
- ROC-AUC
- Ranking quality, threshold-independent
- Confusion matrix
- Per-class error direction, to locate systematic bias
Límites y dificultades
- Under heavy class imbalance, minority classes are ignored while accuracy still looks high
- Wording drift outside the training domain degrades predictions
- Sarcasm, negation and double negation are often flipped, since the literal words point the other way
Conceptos detrás
Aprendizaje supervisado
Pares de pregunta y respuesta enseñan al modelo a responder por sí solo
Embeddings de palabras
Convertir palabras en coordenadas: los sinónimos se agrupan solos y el significado pasa a ser algo que se puede sumar y restar
Evaluación de modelos y validación cruzada
La exactitud es la métrica más fácil de engañar — si evalúas mal, todo lo demás cae