Instructions to use hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking") model = AutoModelForCausalLM.from_pretrained("hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking
- SGLang
How to use hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking with Docker Model Runner:
docker model run hf.co/hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking
CantoneseLLM-v2.0-8B-Thinking
An 8B dense model based on Qwen3-8B that reasons and answers in Hong Kong Cantonese.
It is the final checkpoint of a five-stage pipeline: continual pre-training → chat-vector merge → supervised fine-tuning → DPO → two-stage RLVR. The full development history, including the stages that failed and why, is documented in the technical report.
📄 Paper: CantoneseLLM v2: Reasoning in a Low-Resource Language (arXiv:2609.06970)
🧪 Evaluation Benchmark: hon9kon9ize/hkeval2025
Reasoning in Cantonese
Open models prompted in Cantonese answer in Cantonese but reason in another language. Here is the same translated GSM8K probe, under a system prompt that explicitly asks for Hong Kong Cantonese.
Chat-vector merged checkpoint — reasons entirely in Simplified Chinese, despite the system prompt and despite continual pre-training on Traditional Chinese and Cantonese:
嗯,用户问的是Janet每天赚多少钱。首先,我需要理清楚整个流程。题目说Janet的鸭子每天生16只蛋。
(Well, what the user is asking is how much Janet earns each day. First, I need to work through the whole process. The problem says Janet's ducks lay 16 eggs a day.)
This model — reasons in Cantonese end to end, with the Cantonese classifier 隻 zek3 in place of 只 and colloquial forms such as 剩返 zing6 faan1:
呢個係一個數學應用題,我需要先理解題目入面嘅數據同埋計算邏輯。Janet 隻鴨每日生 16 隻蛋。首先,佢自己食早餐用咗 3 隻,然後用 4 隻嚟整鬆餅。… 剩返嘅蛋就係 16 − 7 = 9 隻。
(This is a mathematics word problem, and I need to understand the figures and the logic of the calculation first. Janet's ducks lay 16 eggs a day. First, she uses 3 for her own breakfast, then 4 to make muffins. … The eggs left over are 16 − 7 = 9.)
Both reach the correct answer, 18 dollars. What changes is where the reasoning happens. The 只/隻 substitution is not cosmetic: 只 can only be read zi2 in Hong Kong Cantonese, which is a different word from the classifier the sentence needs.
The mechanism is a multiplicative language and script term in the RLVR
reward — task_reward × language_multiplier × format_factor — so a
wrong-language trace cannot buy its way back by being correct.
Trace length over eight fixed probes (mean tokens between the reasoning tags):
| Checkpoint | Mean CoT tokens | Median | CoT-to-answer ratio |
|---|---|---|---|
| Qwen3-8B (official) | 491 | 502 | 2.94 |
| Chat-vector merged | 468 | 512 | 3.27 |
| After SFT | 26 | 0 | 0.21 |
| This model | 232 | 237 | 1.09 |
The SFT row is not a typo. At 8B, supervised fine-tuning on a mixture dominated
by another language's reasoning tokens made the model emit an empty reasoning
block on 64.5% of generations — the empty <think></think> pair is a
near-deterministic two-token continuation and became a low-loss attractor. DPO
restored the block; RLVR restored the language.
Benchmark results
HKCanto-Eval (Cheng et al., 2025), reasoning mode on:
| Model | MMLU | CantoMMLU | Cultural | Linguistic | Academic & Prof. | Avg. |
|---|---|---|---|---|---|---|
| Qwen3-8B | 81.08 | 76.68 | 57.94 | 39.50 | 81.61 | 67.36 |
| Chat-vector merged (intermediate) | 80.04 | 76.96 | 66.67 | 41.50 | 82.36 | 69.51 |
| After SFT (intermediate) | 64.23 | 56.15 | 47.62 | 23.00 | 53.97 | 48.99 |
| After DPO (intermediate) | 69.19 | 61.85 | 59.52 | 34.00 | 71.35 | 59.18 |
| This model | 73.86 | 69.73 | 58.33 | 36.00 | 76.08 | 62.80 |
Read this table carefully. SFT cost 20.52 points. DPO recovered 10.19 and RLVR a further 3.62, leaving the model 6.71 points below the merged checkpoint it started from. The 30B-A3B closed the same gap to 1.20 points; this model did not. What it gained instead is Cantonese reasoning, which this benchmark does not measure.
The paper's stated reason is that teaching a model to reason in a new language needs Cantonese reasoning traces present during continual pre-training, not only in post-training — and the 8B had the least headroom to absorb that gap.
If you can run the 30B-A3B, run that one instead. This model is released for deployments where 8B dense is the constraint, and because the contrast between the two sizes is one of the paper's findings.
Artefacts released with this work
Every post-training stage on this card is reproducible from public components. The CPT corpus itself is not released, but its two largest public constituents are.
| What it is | |
|---|---|
| 🏋️ cantonese-nemo-gym-environments | The six NeMo-Gym environments used in RLVR stage 2 — math_with_judge_lang, stem_mcqa, code_gen_lang, workplace_assistant_lang, instruction_following_lang, structured_outputs_lang. Each implements the multiplicative language reward, task_reward × language_multiplier × format_factor. Targets NeMo-Gym 0.3.0rc0; MIT licensed |
| 📊 Nemotron-3-Nano-RL-Training-Blend-STEM-Yue-Translated | The RLVR stage-2 prompt corpus. NVIDIA's Nemotron-3-Nano RL blend translated into Cantonese and Hong Kong Written Chinese, kept as a three-way parallel corpus (en/, yue/, zh-hk/) so the same question can be compared across languages |
| 🌐 Traditional-Chinese-Common-Crawl-by-year | Traditional Chinese extracted from all 111 Common Crawl snapshots, released per snapshot across thirteen years |
| 🇭🇰 Cantonese-Web-Data | The Cantonese subset of the above, filtered with CantoneseDetect and globally deduplicated to 477,298 unique documents |
The three-way parallel structure of the RL corpus is what makes the cross-language pass-rate gap reported in the paper measurable: for the translated environments, prompt language is the only variable.
Other models in this release
All four checkpoints are in the CantoneseLLM v2.0 collection.
| Model | What it is |
|---|---|
| CantoneseLLM-v2.0-8B-Thinking | this model — 8B dense, full pipeline through RLVR |
| CantoneseLLM-v2.0-30B-A3B-Thinking | The flagship. Same five stages at 30B-A3B, and it recovers far more of the SFT regression — prefer it if you can serve it |
| CantoneseLLM-v2.0-8B-Thinking-Chat-Vector-Merged | The intermediate checkpoint this model was built from: CPT + chat-vector merge, before SFT, DPO and RLVR. Scores higher on HKCanto-Eval but reasons in Simplified Chinese — it is the baseline the paper measures against, not a replacement for this model |
| CantoneseLLM-v2.0-30B-A3B-Thinking-Chat-Vector-Merged | The same intermediate stage at 30B-A3B |
Usage
Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
model_id = "hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForCausalLM.from_pretrained(model_id, dtype="auto", device_map="auto")
SYSTEM = "你係CantoneseLLM,一個由Hon9Kon9ize開發嘅語言模型,請使用香港嘅廣東話回答用家問題"
messages = [
{"role": "system", "content": SYSTEM},
{"role": "user", "content": "小明有 5 個蘋果,佢俾咗 2 個朋友,每人 1 個,跟住又買多 3 個。佢而家有幾多個蘋果?"},
]
text = tokenizer.apply_chat_template(messages, tokenize=False, add_generation_prompt=True)
inputs = tokenizer([text], return_tensors="pt").to(model.device)
out = model.generate(**inputs, max_new_tokens=2048, temperature=0.6, top_p=0.95)
print(tokenizer.decode(out[0][len(inputs.input_ids[0]):], skip_special_tokens=True))
vLLM
The stock command works — no special flags are required:
vllm serve hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking
Add --tensor-parallel-size N to shard across N GPUs, and
--reasoning-parser qwen3 if you want the reasoning block returned separately
as reasoning_content rather than inline in content (parser names vary by
vLLM version).
OpenAI-compatible API
The served endpoint speaks the OpenAI protocol, so the official client works unchanged:
from openai import OpenAI
client = OpenAI(base_url="http://localhost:8000/v1", api_key="EMPTY")
SYSTEM = "你係CantoneseLLM,一個由Hon9Kon9ize開發嘅語言模型,請使用香港嘅廣東話回答用家問題"
resp = client.chat.completions.create(
model="hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking",
messages=[
{"role": "system", "content": SYSTEM},
{"role": "user", "content": "小明有 5 個蘋果,佢俾咗 2 個朋友,每人 1 個,跟住又買多 3 個。佢而家有幾多個蘋果?"},
],
temperature=0.6,
top_p=0.95,
max_tokens=2048,
)
msg = resp.choices[0].message
print(getattr(msg, "reasoning_content", None) or "") # populated with --reasoning-parser
print(msg.content)
Or with curl:
curl http://localhost:8000/v1/chat/completions \
-H "Content-Type: application/json" \
-d '{
"model": "hon9kon9ize/CantoneseLLM-v2.0-8B-Thinking",
"messages": [
{"role": "system", "content": "你係CantoneseLLM,一個由Hon9Kon9ize開發嘅語言模型,請使用香港嘅廣東話回答用家問題"},
{"role": "user", "content": "點解香港嘅雨季集中喺五月到九月?"}
],
"temperature": 0.6, "top_p": 0.95, "max_tokens": 2048
}'
Sampling: temperature 0.6, top-p 0.95 — the settings used for every evaluation reported above and in the paper. Avoid greedy decoding.
System prompt: the Cantonese system prompt above was used throughout training and evaluation. Behaviour with other system prompts, or with none, is not characterised.
Thinking and non-thinking. Unlike the 30B-A3B, 25% of this model's SFT data carried no chain-of-thought, deliberately, so that it retains the ability to answer without an explicit reasoning block. It is the closer of the two to stock Qwen3 hybrid behaviour — but given the empty-block history above, verify the mode switch on your own prompts rather than assuming it.
Training pipeline
| Stage | What it installed | Compute |
|---|---|---|
| Continual pre-training | Hong Kong knowledge and Cantonese lexis. 784M tokens, 530 steps, LR 3×10⁻⁵, on 64 TPU v6e chips (Google TRC) via MaxText | 103 TPU chip-hours |
| Chat-vector merge | Instruction following, at no training cost. Δ = Qwen3-8B − Qwen3-8B-Base, added to the CPT checkpoint | — |
| SFT | Translation, data curation and LLM-as-a-judge behaviour. Hybrid mixture, 154.6M tokens over 74,865 rows, 801 steps, LLaMA-Factory | 47 GPU-hours |
| DPO | Repaired the empty-reasoning-block failure. 12,204 preference pairs, 724 steps | 14 GPU-hours |
| RLVR stage 1 | Output format and CoT language, on a multilingual GSM8K blend (50% Cantonese / 25% Written Chinese / 25% English). 150 steps, full fine-tune | 11 GPU-hours |
| RLVR stage 2 | Broadened the reward across six environments. 400 steps, full fine-tune, FP8 end-to-end in both generation and training, with importance-sampling correction | 430 GPU-hours |
Released weights are the final step of stage 2, not a best-validation checkpoint.
Risks & Limitations
- Does not fully recover pre-SFT benchmark performance — 6.71 points below the merged checkpoint. See the table above. The 30B-A3B closed this gap; this model did not.
- Short reasoning traces. 232 tokens on average, at a CoT-to-answer ratio of 1.09. No long original Cantonese reasoning traces existed at the scale SFT needed; human-written or human-verified traces amounted to 131 rows. Most Cantonese CoT in training was machine-translated from English or Simplified Chinese.
- History of empty reasoning blocks. The SFT checkpoint emitted an empty block on 64.5% of generations. DPO and RLVR repaired this, but if you fine-tune further from here, watch for the attractor returning.
- No Hong Kong grounding in the reasoning data. The SFT mixture contained no content grounded in Hong Kong entities or current events — that knowledge comes from CPT only.
- Benchmark scores are underestimates of knowledge. DPO's pair-selection weighted longer responses with no reward for instruction compliance, so the model sometimes ignores "answer with the letter only". Multiple-choice parsers that read a bare leading option letter will under-score it. Use a marker-anchored extractor.
- One run per stage. Severe compute constraints meant no hyperparameter sweep in any post-training stage.
- Preference data were model-judged, not human-annotated, and are bounded by the judge's own Cantonese ability.
- Long context is inherited and untested. Training sequence lengths were 8,192 (SFT/DPO) and 16,384 (RLVR stage 2). Behaviour beyond that is whatever the Qwen3 base provides.
- Standard LLM caveats apply: it will hallucinate, and it has not been safety- tuned beyond what the chat vector carried over.
Citation
@misc{cantonesellm_v2,
title={CantoneseLLM v2: Reasoning in a Low-Resource Language},
author={Tsz Chung Cheng and Chung Shing Cheng and Chaak Ming Lau and Cheuk Hei Chong},
year={2026},
eprint={2609.06970},
archivePrefix={arXiv},
primaryClass={cs.CL},
url={https://arxiv.org/abs/2609.06970},
}
Acknowledgements
Continual pre-training (CPT) was carried out on Cloud TPUs (Tensor Processing Units) from Google's TPU Research Cloud with MaxText. Post-training was carried out on computer resources offered under the category of General Projects by Research Institute for Information Technology, Kyushu University. Usage fee and cost of data-curation costs with proprietary APIs were covered by Votee AI
The RLVR stage builds on NVIDIA's NeMo-RL and NeMo-Gym, and on the Nemotron post-training and RL datasets.
- Downloads last month
- 320