पाठ वर्गीकरण और भावना विश्लेषण
पाठ को पूर्वनिर्धारित लेबल देना
यह पृष्ठ अंग्रेज़ी में प्रस्तुत है; शीर्षक और सारांश का स्थानीयकरण किया गया है।
यह क्षमता क्या है
Takes a piece of text and outputs its category — sentiment polarity, topic, intent or policy violation. The label set is fixed at training time, so it answers which class rather than what is said. A single text can carry several labels, and the model can also emit a probability distribution over the label set.
तकनीकी रूप से कैसे
Early systems paired bag-of-words with linear classifiers; fine-tuning a pre-trained encoder such as the BERT family later became the standard when labelled data is scarce. A classification head sits on top of the encoder and is trained with cross-entropy. With many labels and few examples, it can be recast as sentence-pair entailment or done zero-shot by prompting a large model.
प्रतिनिधि उत्पाद
5GPT-4o
2024मूल रूप से बहुविध सामान्य मॉडल — पाठ, चित्र और ऑडियो एक ही द्वार से
Qwen
2023अनेक आकारों और बहुविध संस्करणों वाला ओपन-वेट परिवार
Hunyuan
2023ओपन-वेट संस्करणों सहित टेनसेंट का सामान्य मॉडल परिवार
ERNIE
2019ज्ञान-संवर्धित प्रीट्रेनिंग से शुरू हुआ चीनी मॉडल, शुरुआती प्रतिनिधि संस्करण
Gemini
2024खोज, ऑफ़िस और बहुविध मॉडल को एक चैट द्वार में समेटता है
संबंधित संस्थान
सामान्य उपयोग
- Sentiment polarity of reviews and public opinion
- Spam and policy-violation filtering
- Routing tickets by topic and priority
- Intent detection to drive dialogue
इसका मूल्यांकन कैसे होता है
- F1
- Harmonic mean of precision and recall; more trustworthy than accuracy under imbalance
- ROC-AUC
- Ranking quality, threshold-independent
- Confusion matrix
- Per-class error direction, to locate systematic bias
सीमाएँ और कठिनाइयाँ
- Under heavy class imbalance, minority classes are ignored while accuracy still looks high
- Wording drift outside the training domain degrades predictions
- Sarcasm, negation and double negation are often flipped, since the literal words point the other way
इसके पीछे की अवधारणाएँ
पर्यवेक्षित अधिगम
‘प्रश्न–उत्तर’ के जोड़े मॉडल को स्वयं उत्तर देना सिखाते हैं
शब्द एम्बेडिंग
शब्दों को निर्देशांक बनाना — समानार्थी स्वयं पास आ जाते हैं और अर्थ को जोड़ना-घटाना संभव हो जाता है
मॉडल मूल्यांकन और क्रॉस-वैलिडेशन
सटीकता सबसे भ्रामक मापदंड है — मूल्यांकन गलत हुआ तो बाकी सब बेकार है