Vietnamese-Tokenizer

Vietnamese-Tokenizer is a 48,000-token Byte-level BPE tokenizer designed primarily for Vietnamese language models trained from scratch.

The tokenizer is optimized for Vietnamese while retaining practical coverage of English and source code. It is intended to be architecture-independent and can be used with Qwen-style, LLaMA-style, or other autoregressive language model architectures as long as the model configuration uses the same vocabulary and token IDs.

Key Features

  • Vocabulary size: 48,000
  • Algorithm: Byte-level BPE
  • Unicode normalization: NFC
  • Lowercasing: No
  • Unknown token: None
  • Fast tokenizer: Yes
  • Byte-level coverage: Any UTF-8 text can be represented
  • Vietnamese-focused vocabulary
  • Includes English and source-code coverage
  • Includes reserved token IDs for future extensions

Special Tokens

Token ID
`< pad
`< bos
`< eos
`< im_start
`< im_end
`< fim_prefix
`< fim_middle
`< fim_suffix

The tokenizer also reserves IDs 8..263 as:

<|reserved_0|>
...
<|reserved_255|>

These reserved tokens are intentionally kept stable so future model variants can introduce additional control tokens without changing the existing token-to-ID mapping.

Training Corpus

The tokenizer was trained on approximately 8 GiB of mixed-domain text.

Approximate composition:

Domain Share
Vietnamese news 70%
Vietnamese Wikipedia 15%
English general text 12%
Source code 3%

The source-code portion includes multiple programming languages, with higher weight assigned to commonly used languages such as Python, JavaScript, Java, C++, and Go.

The corpus used to train this tokenizer is published separately.

Corpus Statistics

The tokenizer training corpus contained approximately:

  • 1,687,254 Vietnamese news documents
  • 1,035,370 Vietnamese Wikipedia documents
  • 326,223 English documents
  • ~107K code documents across multiple programming languages

The total serialized corpus size was approximately 8 GiB.

Usage

from transformers import AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained(
    "PaxiAI/Vietnamese-Tokenizer",
    use_fast=True,
)

text = "Trí tuệ nhân tạo đang thay đổi cách con người làm việc."

ids = tokenizer.encode(
    text,
    add_special_tokens=False,
)

print(ids)
print(tokenizer.decode(ids))

Chat Template

The tokenizer includes a simple ChatML-style template:

<|im_start|>system
You are a helpful assistant.
<|im_end|>
<|im_start|>user
Xin chào!
<|im_end|>
<|im_start|>assistant
Chào bạn!
<|im_end|>

Example:

messages = [
    {
        "role": "system",
        "content": "Bạn là một trợ lý hữu ích."
    },
    {
        "role": "user",
        "content": "Xin chào!"
    }
]

prompt = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
)

print(prompt)

Unicode Handling

The tokenizer applies Unicode NFC normalization.

For example, canonically equivalent NFC and NFD forms of Vietnamese text are normalized to the same representation before tokenization.

This is particularly important for Vietnamese because accented characters may otherwise appear in multiple Unicode representations.

Validation

Before release, the tokenizer passed a production validation suite covering:

  • fixed-text encode/decode round trips
  • Vietnamese NFC/NFD normalization
  • special-token ID stability
  • tokenizer save/reload consistency
  • chat-template rendering
  • randomized Unicode fuzz testing
  • no unknown-token output
  • long-input tokenization
  • fast-tokenizer loading
  • exact vocabulary-size validation

A fuzz test of 10,000 randomly generated Unicode strings completed with zero failures.

The tokenizer also produced zero unknown tokens during validation.

Performance Notes

On manual Vietnamese examples, the tokenizer typically produced roughly 3.6-4.4 characters per token, depending on the sentence.

Example:

Nguyễn Thị Kim Hương đang được cấp cứu tại Bệnh viện Chợ Rẫy.

was tokenized into approximately one token per common Vietnamese syllable or word component.

The tokenizer is intentionally optimized more strongly for Vietnamese than for code identifiers or rare English terms.

Compatibility Warning

Once a model has been pretrained with this tokenizer, the following must remain unchanged:

  • vocabulary
  • token-to-ID mapping
  • normalization rules
  • special-token IDs

Changing any of these creates a different tokenizer and should be released under a new version.

Intended Use

This tokenizer is suitable for:

  • Vietnamese language model pretraining
  • bilingual Vietnamese-English language models
  • Vietnamese instruction-tuned models
  • small and medium autoregressive language models
  • experimental Qwen-style or LLaMA-style architectures
  • models with limited parameter budgets where a very large vocabulary would be inefficient

Limitations

  • The training corpus is heavily weighted toward Vietnamese news and Wikipedia.
  • Conversational Vietnamese is less represented than formal written Vietnamese.
  • Code represents only a small portion of the tokenizer training data.
  • The tokenizer is not intended to be optimal for multilingual models covering many languages equally.
  • A tokenizer alone does not determine model quality; pretraining data and training methodology remain critical.

License

This repository contains the tokenizer artifact itself.

Please ensure that the selected repository license is compatible with the licensing and redistribution requirements of the tokenizer artifact and its training-data sources.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support