🎭 PhoBERT Fine-Tuned for Vietnamese Social Media Emotion Recognition

Mô hình PhoBERT-base được Fine-Tune chuyên biệt cho bài toán nhận diện 7 lớp Cảm xúc tiếng Việt trên dữ liệu Mạng xã hội (Bình luận Facebook, Threads, VNExpress, Teencode/Slang).

Mô hình thuộc hệ thống Microservices dự án SE121 - Social Media with Mental Health Awareness.


📊 7 Lớp Cảm xúc (Emotion Classes)

ID Nhãn Cảm Xúc (English) Tên Tiếng Việt Icon
0 Enjoyment Vui vẻ / Yêu thích 😊
1 Sadness Buồn rầu / Thất vọng 😭
2 Disgust Chán ghét / Khinh bỉ 🤮
3 Anger Tức giận / Bực mình 😡
4 Fear Sợ hãi / Lo lắng 😱
5 Surprise Ngạc nhiên / Bất ngờ 😲
6 Other Khác / Trung tính 😐

📈 Kết quả Đánh giá Thực tế (Version 1.1 Final - Test Set: 1,482 samples)

Emotion Label Precision Recall F1-Score Support
Enjoyment 0.7962 0.7017 0.7459 295
Sadness 0.5887 0.6791 0.6307 215
Disgust 0.5924 0.5203 0.5540 271
Anger 0.7351 0.7083 0.7215 192
Fear 0.6450 0.6301 0.6374 173
Surprise 0.5897 0.6434 0.6154 143
Other 0.5044 0.5907 0.5442 193
Overall Accuracy - - 63.77% 1482
Macro Average 0.6359 0.6391 63.56% 1482
Weighted Average 0.6453 0.6377 63.94% 1482

🚀 Hướng dẫn Sử dụng (Quick Start)

1. Sử dụng Transformers Pipeline

from transformers import pipeline

classifier = pipeline(
    "text-classification",
    model="huyleit/phobert-emotion-social",
    tokenizer="huyleit/phobert-emotion-social",
    top_k=None
)

text = "Bài viết hay quá, đọc mà thấy thích ghê ❤️"
results = classifier(text)
print(results)

2. Sử dụng AutoModel & AutoTokenizer

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

model_name = "huyleit/phobert-emotion-social"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)

text = "Làm có 25-30tr mà cứ tưởng cả tỷ/tháng ko, đòi hỏi vcl ra."
inputs = tokenizer(text, return_tensors="pt")

with torch.no_grad():
    logits = model(**inputs).logits
    probs = torch.softmax(logits, dim=-1)

labels = ["Enjoyment", "Sadness", "Disgust", "Anger", "Fear", "Surprise", "Other"]
for label, prob in zip(labels, probs[0].tolist()):
    print(f"{label}: {prob:.4f}")
Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Evaluation results

  • accuracy on Vietnamese Social Media Emotion Dataset
    self-reported
    0.638
  • f1 on Vietnamese Social Media Emotion Dataset
    self-reported
    0.636