k3-32k tokenizer
Bilingual (Chinese / English) SentencePiece Unigram tokenizer with a 32,000-piece vocabulary, trained for the K3/KQ model family.
Usage
from transformers import AutoTokenizer
tok = AutoTokenizer.from_pretrained("carry121/k3-32k-tokenizer")
ids = tok.encode("深度å¦ä¹ 是机器å¦ä¹ 的一个分支。")
# BOS (<s>, id=2) is prepended automatically (add_bos_token=True).
Both fast (tokenizers backend) and slow (SentencePiece) loading work;
encodings are identical.
Special tokens
| id | piece | role |
|---|---|---|
| 0 | <unk> |
unknown |
| 1 | <pad> |
padding |
| 2 | <s> |
BOS (auto-prepended) |
| 3 | </s> |
EOS |
| — | <zh2en>, <en2zh>, <source>, <target> |
user-defined, atomic (reserved for translation / seq2seq-style tasks) |
IDs 0–3 are fixed and stable; the four user-defined symbols always encode to a single id.
Training
- Algorithm: SentencePiece Unigram, NFKC normalization,
character_coverage=0.9995,byte_fallback=True(256 byte pieces). - Corpus: 2,000,000 deduplicated sentences, sentence-level sampled
- 55% Chinese —
HuggingFaceFW/fineweb-2(cmn_Hani) - 45% English —
HuggingFaceFW/fineweb(sample-10BT)
- 55% Chinese —
- corpus sha256:
4ec6384ada54e71aadc7964f0f0f3aabf0e6d43131bad1766cfcb2ad404448df(dataset revisions are not pinned, so re-training yields a different corpus).
Evaluation (held-out sample sentences)
| language | compression | unk rate | roundtrip |
|---|---|---|---|
| Chinese | 1.70 chars/token | 0 | ✓ |
| English | 5.02 chars/token | 0 | ✓ |
(Reference: an earlier 8K tokenizer scored zh ≈ 1.0, en ≈ 1.09 chars/token.)
Data statement
Trained on text from the FineWeb / FineWeb-2 datasets (ODC-BY). The tokenizer contains no verbatim documents — only subword statistics.
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support