Instructions to use PaxiAI/Vietnamese-Encoder-Base-v1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PaxiAI/Vietnamese-Encoder-Base-v1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("fill-mask", model="PaxiAI/Vietnamese-Encoder-Base-v1")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("PaxiAI/Vietnamese-Encoder-Base-v1") model = AutoModel.from_pretrained("PaxiAI/Vietnamese-Encoder-Base-v1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
PaxiAI Vietnamese Encoder Base v1
PaxiAI Vietnamese Encoder Base v1 is a Vietnamese-focused bidirectional Transformer encoder pretrained from scratch using Masked Language Modeling (MLM). It uses the architecture principles of ModernBERT, all model parameters trained from scratch with the PaxiAI Vietnamese Tokenizer.
This model is intended to serve as a reusable foundation encoder for Vietnamese NLP tasks such as:
- semantic embeddings
- information retrieval
- reranking
- text classification
- intent detection
- named entity recognition
- semantic similarity
- sentence representation
- domain-specific encoder models
Model Overview
| Property | Value |
|---|---|
| Architecture | Bidirectional Transformer Encoder |
| Training objective | Masked Language Modeling |
| Initialization | Random weights |
| Hidden size | 768 |
| Transformer layers | 12 |
| Attention heads | 12 |
| Intermediate size | 2048 |
| Vocabulary size | 48,000 |
| Maximum sequence length | 2,048 |
| Primary language | Vietnamese |
| Tokenizer | PaxiAI/Vietnamese-Tokenizer |
| Pretraining steps | 200,000 |
| Pretrained model dependency | None |
Tokenizer
The model uses:
PaxiAI/Vietnamese-Tokenizer
https://huggingface.co/PaxiAI/Vietnamese-Tokenizer
The tokenizer has a vocabulary size of approximately 48K tokens and was designed primarily for Vietnamese text while retaining support for English and programming-related text.
For MLM pretraining, one of the tokenizer's reserved tokens is used as the mask token without modifying the original vocabulary or token ID mapping.
Loading the Model
from transformers import AutoTokenizer, AutoModel
model_id = "PaxiAI/Vietnamese-Encoder-Base-v1"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModel.from_pretrained(model_id)
text = "Trí tuệ nhân tạo đang phát triển rất nhanh."
inputs = tokenizer(
text,
return_tensors="pt",
truncation=True,
max_length=512,
)
outputs = model(**inputs)
print(outputs.last_hidden_state.shape)
The output contains contextual token representations:
[batch_size, sequence_length, 768]
License
Please refer to the license associated with this repository and verify the licenses and usage conditions of the individual datasets used during pretraining.
- Downloads last month
- 5