PhoBERT-VSMEC-Emotion (Single-Label)

Mô hình phân loại cảm xúc tiếng Việt (Emotion Classification) được fine-tune từ vinai/phobert-base-v2 trên bộ dữ liệu UIT-VSMEC (Vietnamese Students’ Multilabel Emotion Corpus).

Mô hình này được tối ưu hóa để nhận diện 01 cảm xúc chủ đạo (Single-label) và có khả năng hiểu ngữ nghĩa của các Icon/Emoji phổ biến.

Thông tin mô hình

  • Base Model: vinai/phobert-base-v2
  • Task: Single-label Emotion Classification (Phân loại đơn nhãn),
  • Số lượng nhãn (7 classes): Other, Disgust, Enjoyment, Sadness, Fear, Surprise, Anger.
  • Tách từ (Segmentation): Sử dụng VnCoreNLP (annotator: wseg).
  • Tính năng đặc biệt (Emoji Awareness):
    • Tokenizer được thêm 15 icon/emoji phổ biến (ví dụ: 😡, 😭, 😂, 👍, 😱...).
    • Các token emoji này được khởi tạo embedding vector từ các từ tiếng Việt tương ứng (ví dụ: vector "😡" được copy từ vector "giận", "😭" từ "khóc") giúp mô hình hiểu cảm xúc qua icon tốt hơn ngay từ đầu.

Cấu hình huấn luyện

  • Epochs: 10
  • Batch size: 16 (train) / 32 (eval)
  • Learning rate: 2e-5
  • Metric tối ưu: F1-Macro
  • Precision: FP16

Kết quả đánh giá (Test Set)

Kết quả đánh giá trên tập test của UIT-VSMEC (Single-label metrics):

Metric Score
Accuracy 0.6580
F1-Macro 0.6291
F1-Weighted 0.6572

(Nguồn: Cell trong notebook)

Hướng dẫn sử dụng (Inference)

from transformers import AutoTokenizer, AutoModelForSequenceClassification
from vncorenlp import VnCoreNLP
import torch
import numpy as np

# 1. Load Model & Tokenizer
model_name = "quyle1304/PhoBERT-EmotionClassifier" # Thay bằng đường dẫn model của bạn
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSequenceClassification.from_pretrained(model_name)

# 2. Setup VnCoreNLP (Bắt buộc để tách từ)
# Tải VnCoreNLP-1.1.1.jar và cam/ tại: https://github.com/vncorenlp/VnCoreNLP
rdr = VnCoreNLP("path/to/VnCoreNLP-1.1.1.jar", annotators="wseg") 

def segment(text):
    return " ".join([" ".join(sent) for sent in rdr.tokenize(text)])

# 3. Predict Function
def predict(text):
    text_seg = segment(text)
    inputs = tokenizer(text_seg, return_tensors="pt", truncation=True, max_length=256)
    
    with torch.no_grad():
        outputs = model(**inputs)
        logits = outputs.logits
        probs = torch.softmax(logits, dim=1).cpu().numpy()[0] # Dùng Softmax cho Single-label
        
    labels = ["Other", "Disgust", "Enjoyment", "Sadness", "Fear", "Surprise", "Anger"]
    pred_label = labels[np.argmax(probs)]
    confidence = np.max(probs)
    
    return pred_label, confidence

# 4. Run
examples = [
    "Hôm nay trời đẹp quá, mình cảm thấy rất vui! 😂",
    "Phim này chán ngắt, phí cả tiền vé 😡"
]

for text in examples:
    label, conf = predict(text)
    print(f"Text: {text} | Emotion: {label} ({conf:.2%})")
Downloads last month
13
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support