DistilPhoBERT

DistilPhoBERT is a distilled version of PhoBERT-base designed for Vietnamese Natural Language Processing (NLP). The model is trained using knowledge distillation to significantly reduce model size and inference latency while preserving the performance of the original PhoBERT model across multiple downstream tasks.

  • Base model: PhoBERT-base
  • Language: Vietnamese
  • Architecture: RoBERTa
  • Training objective: Masked Language Modeling (MLM)

Model Details

Property Value
Base architecture RoBERTa
Teacher model PhoBERT-base
Hidden size 768
Attention heads 12
Transformer layers 6
Parameters 92M
Model size 352 MB

The student model is initialized using skip-layer mapping, where the 6 Transformer layers are copied from layers 0, 2, 4, 6, 8, and 10 of the teacher model before knowledge distillation.


Key Features

  • 32% fewer parameters

    • PhoBERT-base: 135M
    • DistilPhoBERT: 92M
  • 1.88×–1.90× faster inference

  • Excellent performance preservation

    • Retains 96.3% – 99.8% of PhoBERT's performance on the VLSP 2016 benchmark.
  • Pre-trained on approximately 20GB of Vietnamese news articles


Training

DistilPhoBERT is trained using a teacher–student knowledge distillation framework.

The optimization objective combines three complementary losses:

  • Masked Language Modeling (MLM) Loss

    • Enables the student model to learn contextual representations by predicting masked tokens.
  • Knowledge Distillation (KD) Loss

    • Uses KL Divergence with a temperature of 2.0 to align the output probability distributions of the student and teacher models.
  • Cosine Embedding Loss

    • Encourages the hidden representations of the student to match those of the teacher.

Training Data

The model is pre-trained on a cleaned Vietnamese news corpus derived from: ademax/binhvq-news-corpus

The preprocessing pipeline includes:

  • HTML and boilerplate removal
  • Unicode NFC normalization
  • Punctuation normalization
  • Author and signature removal
  • Quality filtering
  • Duplicate removal using MD5 hashing
  • Vietnamese word segmentation using VnCoreNLP

Processed Dataset

The processed dataset is publicly available on Hugging Face: trungbb8/vietnamese-news-copus-segmented


Evaluation

The model was fine-tuned and evaluated on the VLSP 2016 benchmark.

Task PhoBERT-base DistilPhoBERT Performance Retained
POS Tagging 0.9486 0.9472 99.85%
Named Entity Recognition 0.9285 0.9246 99.57%
Sentiment Analysis 0.7790 0.7505 96.34%

Overall, DistilPhoBERT achieves nearly the same performance as PhoBERT while reducing inference latency by approximately 1.9×.


Usage

Load the model

from transformers import AutoTokenizer, AutoModel

model_name = "trungbb8/distilphobert"

tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModel.from_pretrained(model_name)

# Example sentence (Must be word-segmented)
text = "Trường Đại_học Nông_Lâm Thành_phố Hồ_Chí_Minh."

inputs = tokenizer(text, return_tensors="pt")
outputs = model(**inputs)

last_hidden_state = outputs.last_hidden_state

Fine-tuning

For downstream tasks, use the corresponding Hugging Face model classes.

Text Classification

from transformers import AutoModelForSequenceClassification

Named Entity Recognition

from transformers import AutoModelForTokenClassification

Intended Uses

DistilPhoBERT can be used as the backbone model for a wide range of Vietnamese NLP applications, including:

  • Text Classification
  • Sentiment Analysis
  • Named Entity Recognition (NER)
  • POS Tagging
  • Semantic Search
  • Information Retrieval
  • Sentence Embedding
  • Retrieval-Augmented Generation (RAG)

Limitations

  • The model is trained exclusively on Vietnamese text.
  • Performance may decrease on specialized domains that differ significantly from the news corpus.
  • The model is intended primarily for research and educational purposes. Users should evaluate its performance before deploying it in production environments.

Citation

@misc{distilphobert2026,
  title={DistilPhoBERT: A Compressed Variant of PhoBERT for Vietnamese Natural Language Processing},
  author={Nguyen Minh Trung and Nguyen Thanh Thuong},
  year={2026},
  note={Bachelor Thesis, Faculty of Information Technology, Nong Lam University},
  url={https://huggingface.co/trungbb8/distilphobert}
}

Downloads last month
22
Safetensors
Model size
92.5M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Datasets used to train trungbb8/distilphobert