Qwen3.5-9B-Coding-MixedOPD-TMAX27B
One training run, multiple full-weight checkpoints on branches. This revision contains u109.
默认 main 为 u109;每个 uN 分支保存对应 checkpoint 的完整权重、训练来源和评测记录。
Checkpoint branches and orch evaluations
| Branch | Orch eval150 resolved | Evaluation resource profile |
|---|---|---|
u109 / main |
87/150 (58.0%) | CPU-quota thread limits; gcp2-w32-cpuquota-20260920 |
| u307 | 84/150 (56.0%) | CPU-quota thread limits; gcp2-w32-cpuquota-20260920 |
| u347 | 84/150 (56.0%) | CPU-quota thread limits; gcp1 |
main uses the strongest result within the current matched-control cohort (ties prefer the later checkpoint).
Historical site/thread settings are labeled above; their scores are separate observations, not pooled comparisons.
The primary evaluation is orch: MiniMax-M2.7 coordinates and this checkpoint is the small coding worker.
This branch: 87/150; matching actual starting-model control: 81/150, difference +6 tasks. Paired outcomes and control fingerprint are included in EVAL_RESULTS.json.
All listed primary evaluations completed 150 tasks, passed their evaluator audits, and had zero final infrastructure errors. Model/test failures and verifier timeouts count as unresolved. Every selected checkpoint's available completed historical repetitions are preserved in EVAL_RESULTS.json. Selection used eval150 results, so these are exploratory candidates, not independent confirmation of a gain or an official full SWE-bench leaderboard score.
Training run and provenance
- W&B run: 9564fce0
- Run name:
resume-step109plus4_mix75new25old_tmax27b_b64g4_lr5e-7_to-u110-opd - Campaign:
opd-step109-newdata-20260915 - Lineage: Original HF OPD step109 → 4 new-data OPD updates → mixed-data OPD with TMAX-27B.
- Training algorithm / mode: On-policy distillation (OPD) /
orch. - Training data: 75% new verified SWE-rebench tasks + 25% historical SWE-bench train350, by prompt groups per update.
- Local checkpoint: u109. uN is the recorded local checkpoint iteration label, not the original HF OPD step number. Solo350 continues numbering from the earlier local63 checkpoint.
- Optimizer: weights-only resumes use a fresh optimizer; this run includes restarts.
Teacher: allenai/tmax-27b; MiniMax-M2.7 coordinates training work orders. OPD teacher KL coefficient is 1.0; reference-policy KL is 0. Source pool: 774 accepted new tasks plus 350 historical tasks, sampled 75/25 by prompt groups. Runtime image exclusions are recorded separately in TRAINING_CONFIG.json.
| Training setting | Value |
|---|---|
| Learning rate | 5e-07 |
| Samples per update | 64 |
| Group size | 4 |
| Prompt groups per update | 16 |
| Reference-policy KL coefficient | 0.0 |
| Maximum weight staleness | 1 |
| Save interval | 4 updates |
| Worker thinking | disabled |
Exact parent revisions, dataset hashes, archived training settings and native checkpoint identity are in CHECKPOINT_PROVENANCE.json and TRAINING_CONFIG.json.
Evaluation settings
150 held-out tasks, 32 episodes per job, 2 CPUs and 10 GiB per sandbox; worker thinking disabled. Worker/coordinator temperatures: 0.2; maximum worker turns: 64; response budgets: 8192/4096 tokens; verifier timeout: 2400 seconds. CPU-quota profiles additionally set Django/numeric thread defaults to 2. Exact per-branch settings, task outcomes, fingerprints, repetitions and resource profiles are in EVAL_RESULTS.json.
Download a checkpoint
from huggingface_hub import snapshot_download
path = snapshot_download("CharlieLLL/Qwen3.5-9B-Coding-MixedOPD-TMAX27B", revision="u109")
Use a serving stack with Qwen3.5 support and the included tokenizer/chat template. These are full BF16 weights plus tokenizer/processor assets; optimizer state is not included. Evaluation covers text-based coding tasks; inherited vision configuration does not establish evaluated multimodal behavior. License: Apache-2.0; see LICENSE.
- Downloads last month
- 24
Model tree for CharlieLLL/Qwen3.5-9B-Coding-MixedOPD-TMAX27B
Base model
Shangy/Qwen3.5-9B-OPD-Coding