Phân loại văn bản và phân tích cảm xúc
Gán nhãn định trước cho một đoạn văn bản
Toàn văn của trang này được trình bày bằng tiếng Anh; tiêu đề và phần dẫn đã được bản địa hóa.
NĂNG LỰC NÀY NGHĨA LÀ GÌ
Takes a piece of text and outputs its category — sentiment polarity, topic, intent or policy violation. The label set is fixed at training time, so it answers which class rather than what is said. A single text can carry several labels, and the model can also emit a probability distribution over the label set.
Làm ra sao về mặt kỹ thuật
Early systems paired bag-of-words with linear classifiers; fine-tuning a pre-trained encoder such as the BERT family later became the standard when labelled data is scarce. A classification head sits on top of the encoder and is trained with cross-entropy. With many labels and few examples, it can be recast as sentence-pair entailment or done zero-shot by prompting a large model.
Sản phẩm tiêu biểu
5GPT-4o
2024Mô hình đa phương thức gốc, xử lý văn bản, hình ảnh và âm thanh qua một cửa vào
Qwen
2023Dòng trọng số mở với nhiều kích cỡ và bản đa phương thức
Hunyuan
2023Dòng mô hình phổ thông của Tencent, có bản trọng số mở
ERNIE
2019Mô hình tiếng Trung khởi đầu bằng tiền huấn luyện tăng cường tri thức, bản đầu tiên tiêu biểu
Gemini
2024Một cửa vào trò chuyện gom tìm kiếm, ứng dụng văn phòng và mô hình đa phương thức
Tổ chức liên quan
Cách dùng tiêu biểu
- Sentiment polarity of reviews and public opinion
- Spam and policy-violation filtering
- Routing tickets by topic and priority
- Intent detection to drive dialogue
Đánh giá nó tốt hay không thế nào
- F1
- Harmonic mean of precision and recall; more trustworthy than accuracy under imbalance
- ROC-AUC
- Ranking quality, threshold-independent
- Confusion matrix
- Per-class error direction, to locate systematic bias
Ranh giới và điểm khó
- Under heavy class imbalance, minority classes are ignored while accuracy still looks high
- Wording drift outside the training domain degrades predictions
- Sarcasm, negation and double negation are often flipped, since the literal words point the other way
Các khái niệm đằng sau
Học có giám sát
Từng cặp câu hỏi – đáp án dạy mô hình tự trả lời
Nhúng từ (Word Embeddings)
Biến từ thành tọa độ: từ đồng nghĩa tự tụ lại và ý nghĩa lần đầu có thể cộng trừ
Đánh giá mô hình và kiểm định chéo
Độ chính xác là chỉ số dễ đánh lừa nhất — đánh giá sai thì mọi thứ sau đó đều vô nghĩa