UD_model_v1

Full-parameter SFT of Qwen3Guard-Gen-4B for Hong Kong PDPO Data Protection Principle classification, continued from the project's earlier PDPO fine-tuned checkpoint.

Training

The continued stage used 2,000 mixed Mandarin/Hong Kong Chinese/English conversations, with 64 validation examples and 600 test examples. It trained for two epochs with learning rate 5e-6, effective batch size 16, BF16 compute and the LLaMA-Factory qwen3_nothink template. The optimizer was initialized afresh. Weights are saved in FP32; serving can load them in BF16. This repository contains full model weights, not a LoRA adapter.

Pair IDs and identical normalized conversation groups do not cross the new splits. Test pairs and text were excluded from the earlier prepared SFT data. The source dataset was evaluated in prior experiments, so the holdout is not a never-inspected external benchmark.

Evaluation

Checkpoint Accuracy Macro F1
Earlier PDPO checkpoint 83.00% 71.28%
UD_model_v1 95.33% 93.55%

Both checkpoints were evaluated on the same 600 examples with the training-compatible template. Raw and normalized label accuracy were identical; there were zero invalid outputs and request errors. These results measure performance against synthetic/source annotations, not legal correctness or general safety across arbitrary domains.

Important: template and prompt

This release packages the corrected chat_template.jinja, preserving system/user roles and the training-compatible assistant prefix. The local training output originally retained Qwen3Guard's native safety wrapper; that wrapper injected conflicting Safe/Unsafe/Controversial instructions. It is not used in this release. Model weights are unchanged by this packaging correction.

Use the DPP system prompt and input formatting supplied by the repository evaluator. The model returns JSON fields label, risk_level, violated_rule_ids and explanation. Labels are DPP1–DPP6 or NONE. Risk labels are HIGH/LOW proxies, not independently annotated severity.

Example serving command:

vllm serve JeremyLiShuhao/UD_model_v1 \
  --served-model-name ud-model-v1 \
  --host 127.0.0.1 --port 8015 --dtype bfloat16 \
  --max-model-len 8192 --generation-config vllm

See the benchmark repository for explicit --chat-template usage and the full evaluation command. No label-enum constrained decoding was used for the reported accuracy.

Limitations

This is a research classifier, not legal advice or a replacement for human review. Performance may change on other languages, longer conversations, adversarial inputs and real-world distributions. Follow the upstream model's applicable license and terms; this card does not grant additional rights to source datasets.

Downloads last month
260
Safetensors
Model size
4B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JeremyLiShuhao/UD_model_v1

Finetuned
Qwen/Qwen3-4B
Finetuned
(7)
this model