🎭 PhoBERT Fine-Tuned for Vietnamese Social Media Emotion Recognition
Mô hình PhoBERT-base được Fine-Tune chuyên biệt cho bài toán nhận diện 7 lớp Cảm xúc tiếng Việt trên dữ liệu Mạng xã hội (Bình luận Facebook, Threads, VNExpress, Teencode/Slang).
Mô hình thuộc hệ thống Microservices dự án SE121 - Social Media with Mental Health Awareness.
📊 7 Lớp Cảm xúc (Emotion Classes)
| ID | Nhãn Cảm Xúc (English) | Tên Tiếng Việt | Icon |
|---|---|---|---|
| 0 | Enjoyment | Vui vẻ / Yêu thích | 😊 |
| 1 | Sadness | Buồn rầu / Thất vọng | 😭 |
| 2 | Disgust | Chán ghét / Khinh bỉ | 🤮 |
| 3 | Anger | Tức giận / Bực mình | 😡 |
| 4 | Fear | Sợ hãi / Lo lắng | 😱 |
| 5 | Surprise | Ngạc nhiên / Bất ngờ | 😲 |
| 6 | Other | Khác / Trung tính | 😐 |
📈 Kết quả Đánh giá Thực tế (Version 1.1 Final - Test Set: 1,482 samples)
| Emotion Label | Precision | Recall | F1-Score | Support |
|---|---|---|---|---|
| Enjoyment | 0.7962 | 0.7017 | 0.7459 | 295 |
| Sadness | 0.5887 | 0.6791 | 0.6307 | 215 |
| Disgust | 0.5924 | 0.5203 | 0.5540 | 271 |
| Anger | 0.7351 | 0.7083 | 0.7215 | 192 |
| Fear | 0.6450 | 0.6301 | 0.6374 | 173 |
| Surprise | 0.5897 | 0.6434 | 0.6154 | 143 |
| Other | 0.5044 | 0.5907 | 0.5442 | 193 |
| Overall Accuracy | - | - | 63.77% | 1482 |
| Macro Average | 0.6359 | 0.6391 | 63.56% | 1482 |
| Weighted Average | 0.6453 | 0.6377 | 63.94% | 1482 |
🚀 Hướng dẫn Sử dụng (Quick Start)
1. Sử dụng Transformers Pipeline
from transformers import pipeline
classifier = pipeline(
"text-classification",
model="huyleit/phobert-emotion-social",
tokenizer="huyleit/phobert-emotion-social",
top_k=None
)
text = "Bài viết hay quá, đọc mà thấy thích ghê ❤️"
results = classifier(text)
print(results)
2. Sử dụng AutoModel & AutoTokenizer
import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification
model_name = "huyleit/phobert-emotion-social"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)
text = "Làm có 25-30tr mà cứ tưởng cả tỷ/tháng ko, đòi hỏi vcl ra."
inputs = tokenizer(text, return_tensors="pt")
with torch.no_grad():
logits = model(**inputs).logits
probs = torch.softmax(logits, dim=-1)
labels = ["Enjoyment", "Sadness", "Disgust", "Anger", "Fear", "Surprise", "Other"]
for label, prob in zip(labels, probs[0].tolist()):
print(f"{label}: {prob:.4f}")
- Downloads last month
- -
Evaluation results
- accuracy on Vietnamese Social Media Emotion Datasetself-reported0.638
- f1 on Vietnamese Social Media Emotion Datasetself-reported0.636