violetxi/qwen35-9b-wmrl-v4-R-only
wm-internalization v4 checkpoint — condition R-only, save final.
Full-finetune of Qwen/Qwen3.5-9B on the Calderwood & Harkness synthetic
law-firm corpus (world-internalization study, v4 lineage: 9B student,
~50k think-on seed pool). Grafted back into the hub composite layout
(Qwen3_5ForConditionalGeneration) — servable with vLLM out of the box.
- training data: see
train_summary.jsonin the training run directory - graft: {"trained": "/scratch/11457/ziyxiang/wm-rl-runs/ckpts-v4/R-only/final", "ref": "/scratch/11457/ziyxiang/.cache/huggingface/hub/models--Qwen--Qwen3.5-9B/snapshots/c202236235762e1c871ad0ccb60c8ee5ba337b9a", "replaced": 427}
- uploaded: 2026-09-07T14:18:57+00:00 by hf_upload.py (PLAN4.md F-D policy)
- Downloads last month
- -
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support