k3-32k tokenizer

Bilingual (Chinese / English) SentencePiece Unigram tokenizer with a 32,000-piece vocabulary, trained for the K3/KQ model family.

Usage

from transformers import AutoTokenizer

tok = AutoTokenizer.from_pretrained("carry121/k3-32k-tokenizer")
ids = tok.encode("深度学习是机器学习的一个分支。")
# BOS (<s>, id=2) is prepended automatically (add_bos_token=True).

Both fast (tokenizers backend) and slow (SentencePiece) loading work; encodings are identical.

Special tokens

id piece role
0 <unk> unknown
1 <pad> padding
2 <s> BOS (auto-prepended)
3 </s> EOS
— <zh2en>, <en2zh>, <source>, <target> user-defined, atomic (reserved for translation / seq2seq-style tasks)

IDs 0–3 are fixed and stable; the four user-defined symbols always encode to a single id.

Training

  • Algorithm: SentencePiece Unigram, NFKC normalization, character_coverage=0.9995, byte_fallback=True (256 byte pieces).
  • Corpus: 2,000,000 deduplicated sentences, sentence-level sampled
    • 55% Chinese — HuggingFaceFW/fineweb-2 (cmn_Hani)
    • 45% English — HuggingFaceFW/fineweb (sample-10BT)
  • corpus sha256: 4ec6384ada54e71aadc7964f0f0f3aabf0e6d43131bad1766cfcb2ad404448df (dataset revisions are not pinned, so re-training yields a different corpus).

Evaluation (held-out sample sentences)

language compression unk rate roundtrip
Chinese 1.70 chars/token 0 ✓
English 5.02 chars/token 0 ✓

(Reference: an earlier 8K tokenizer scored zh ≈ 1.0, en ≈ 1.09 chars/token.)

Data statement

Trained on text from the FineWeb / FineWeb-2 datasets (ODC-BY). The tokenizer contains no verbatim documents — only subword statistics.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support