텍스트 분류와 감정 분석
글에 미리 정한 라벨을 붙인다
이 페이지의 본문은 영어로 제공됩니다. 제목과 요약은 한국어로 번역되었습니다.
이 능력이 뜻하는 것
Takes a piece of text and outputs its category — sentiment polarity, topic, intent or policy violation. The label set is fixed at training time, so it answers which class rather than what is said. A single text can carry several labels, and the model can also emit a probability distribution over the label set.
기술적으로 구현하는 방법
Early systems paired bag-of-words with linear classifiers; fine-tuning a pre-trained encoder such as the BERT family later became the standard when labelled data is scarce. A classification head sits on top of the encoder and is trained with cross-entropy. With many labels and few examples, it can be recast as sentence-pair entailment or done zero-shot by prompting a large model.
대표 제품
5GPT-4o
2024네이티브 멀티모달 범용 모델. 텍스트·이미지·오디오를 한 창구에서 다룬다
Qwen
2023여러 규모와 멀티모달 버전을 아우르는 오픈웨이트 모델 계열
Hunyuan
2023오픈웨이트 버전을 포함한 텐센트의 범용 모델 계열
ERNIE
2019지식 강화 사전학습에서 출발한 중국어 모델, 초기 대표 버전
Gemini
2024검색·오피스·멀티모달 모델을 하나의 대화 창구로 모았다
관련 기관
대표적 용도
- Sentiment polarity of reviews and public opinion
- Spam and policy-violation filtering
- Routing tickets by topic and priority
- Intent detection to drive dialogue
성능을 평가하는 방법
- F1
- Harmonic mean of precision and recall; more trustworthy than accuracy under imbalance
- ROC-AUC
- Ranking quality, threshold-independent
- Confusion matrix
- Per-class error direction, to locate systematic bias
경계와 난점
- Under heavy class imbalance, minority classes are ignored while accuracy still looks high
- Wording drift outside the training domain degrades predictions
- Sarcasm, negation and double negation are often flipped, since the literal words point the other way