Instructions to use PaxiAI/Vietnamese-Tokenizer with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use PaxiAI/Vietnamese-Tokenizer with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="PaxiAI/Vietnamese-Tokenizer") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoModel model = AutoModel.from_pretrained("PaxiAI/Vietnamese-Tokenizer", device_map="auto") - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use PaxiAI/Vietnamese-Tokenizer with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "PaxiAI/Vietnamese-Tokenizer" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PaxiAI/Vietnamese-Tokenizer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/PaxiAI/Vietnamese-Tokenizer
- SGLang
How to use PaxiAI/Vietnamese-Tokenizer with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "PaxiAI/Vietnamese-Tokenizer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PaxiAI/Vietnamese-Tokenizer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "PaxiAI/Vietnamese-Tokenizer" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "PaxiAI/Vietnamese-Tokenizer", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use PaxiAI/Vietnamese-Tokenizer with Docker Model Runner:
docker model run hf.co/PaxiAI/Vietnamese-Tokenizer
Vietnamese-Tokenizer
Vietnamese-Tokenizer is a 48,000-token Byte-level BPE tokenizer designed primarily for Vietnamese language models trained from scratch.
The tokenizer is optimized for Vietnamese while retaining practical coverage of English and source code. It is intended to be architecture-independent and can be used with Qwen-style, LLaMA-style, or other autoregressive language model architectures as long as the model configuration uses the same vocabulary and token IDs.
Key Features
- Vocabulary size: 48,000
- Algorithm: Byte-level BPE
- Unicode normalization: NFC
- Lowercasing: No
- Unknown token: None
- Fast tokenizer: Yes
- Byte-level coverage: Any UTF-8 text can be represented
- Vietnamese-focused vocabulary
- Includes English and source-code coverage
- Includes reserved token IDs for future extensions
Special Tokens
| Token | ID |
|---|---|
| `< | pad |
| `< | bos |
| `< | eos |
| `< | im_start |
| `< | im_end |
| `< | fim_prefix |
| `< | fim_middle |
| `< | fim_suffix |
The tokenizer also reserves IDs 8..263 as:
<|reserved_0|>
...
<|reserved_255|>
These reserved tokens are intentionally kept stable so future model variants can introduce additional control tokens without changing the existing token-to-ID mapping.
Training Corpus
The tokenizer was trained on approximately 8 GiB of mixed-domain text.
Approximate composition:
| Domain | Share |
|---|---|
| Vietnamese news | 70% |
| Vietnamese Wikipedia | 15% |
| English general text | 12% |
| Source code | 3% |
The source-code portion includes multiple programming languages, with higher weight assigned to commonly used languages such as Python, JavaScript, Java, C++, and Go.
The corpus used to train this tokenizer is published separately.
Corpus Statistics
The tokenizer training corpus contained approximately:
- 1,687,254 Vietnamese news documents
- 1,035,370 Vietnamese Wikipedia documents
- 326,223 English documents
- ~107K code documents across multiple programming languages
The total serialized corpus size was approximately 8 GiB.
Usage
from transformers import AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained(
"PaxiAI/Vietnamese-Tokenizer",
use_fast=True,
)
text = "Trí tuệ nhân tạo đang thay đổi cách con người làm việc."
ids = tokenizer.encode(
text,
add_special_tokens=False,
)
print(ids)
print(tokenizer.decode(ids))
Chat Template
The tokenizer includes a simple ChatML-style template:
<|im_start|>system
You are a helpful assistant.
<|im_end|>
<|im_start|>user
Xin chào!
<|im_end|>
<|im_start|>assistant
Chào bạn!
<|im_end|>
Example:
messages = [
{
"role": "system",
"content": "Bạn là một trợ lý hữu ích."
},
{
"role": "user",
"content": "Xin chào!"
}
]
prompt = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
)
print(prompt)
Unicode Handling
The tokenizer applies Unicode NFC normalization.
For example, canonically equivalent NFC and NFD forms of Vietnamese text are normalized to the same representation before tokenization.
This is particularly important for Vietnamese because accented characters may otherwise appear in multiple Unicode representations.
Validation
Before release, the tokenizer passed a production validation suite covering:
- fixed-text encode/decode round trips
- Vietnamese NFC/NFD normalization
- special-token ID stability
- tokenizer save/reload consistency
- chat-template rendering
- randomized Unicode fuzz testing
- no unknown-token output
- long-input tokenization
- fast-tokenizer loading
- exact vocabulary-size validation
A fuzz test of 10,000 randomly generated Unicode strings completed with zero failures.
The tokenizer also produced zero unknown tokens during validation.
Performance Notes
On manual Vietnamese examples, the tokenizer typically produced roughly 3.6-4.4 characters per token, depending on the sentence.
Example:
Nguyễn Thị Kim Hương đang được cấp cứu tại Bệnh viện Chợ Rẫy.
was tokenized into approximately one token per common Vietnamese syllable or word component.
The tokenizer is intentionally optimized more strongly for Vietnamese than for code identifiers or rare English terms.
Compatibility Warning
Once a model has been pretrained with this tokenizer, the following must remain unchanged:
- vocabulary
- token-to-ID mapping
- normalization rules
- special-token IDs
Changing any of these creates a different tokenizer and should be released under a new version.
Intended Use
This tokenizer is suitable for:
- Vietnamese language model pretraining
- bilingual Vietnamese-English language models
- Vietnamese instruction-tuned models
- small and medium autoregressive language models
- experimental Qwen-style or LLaMA-style architectures
- models with limited parameter budgets where a very large vocabulary would be inefficient
Limitations
- The training corpus is heavily weighted toward Vietnamese news and Wikipedia.
- Conversational Vietnamese is less represented than formal written Vietnamese.
- Code represents only a small portion of the tokenizer training data.
- The tokenizer is not intended to be optimal for multilingual models covering many languages equally.
- A tokenizer alone does not determine model quality; pretraining data and training methodology remain critical.
License
This repository contains the tokenizer artifact itself.
Please ensure that the selected repository license is compatible with the licensing and redistribution requirements of the tokenizer artifact and its training-data sources.