Qwen3.5-9B-Coding-MixedOPD-TMAX27B

One training run, multiple full-weight checkpoints on branches. This revision contains u109. 默认 main 为 u109;每个 uN 分支保存对应 checkpoint 的完整权重、训练来源和评测记录。

Checkpoint branches and orch evaluations

Branch Orch eval150 resolved Evaluation resource profile
u109 / main 87/150 (58.0%) CPU-quota thread limits; gcp2-w32-cpuquota-20260920
u307 84/150 (56.0%) CPU-quota thread limits; gcp2-w32-cpuquota-20260920
u347 84/150 (56.0%) CPU-quota thread limits; gcp1

main uses the strongest result within the current matched-control cohort (ties prefer the later checkpoint). Historical site/thread settings are labeled above; their scores are separate observations, not pooled comparisons. The primary evaluation is orch: MiniMax-M2.7 coordinates and this checkpoint is the small coding worker.

This branch: 87/150; matching actual starting-model control: 81/150, difference +6 tasks. Paired outcomes and control fingerprint are included in EVAL_RESULTS.json.

All listed primary evaluations completed 150 tasks, passed their evaluator audits, and had zero final infrastructure errors. Model/test failures and verifier timeouts count as unresolved. Every selected checkpoint's available completed historical repetitions are preserved in EVAL_RESULTS.json. Selection used eval150 results, so these are exploratory candidates, not independent confirmation of a gain or an official full SWE-bench leaderboard score.

Training run and provenance

  • W&B run: 9564fce0
  • Run name: resume-step109plus4_mix75new25old_tmax27b_b64g4_lr5e-7_to-u110-opd
  • Campaign: opd-step109-newdata-20260915
  • Lineage: Original HF OPD step109 → 4 new-data OPD updates → mixed-data OPD with TMAX-27B.
  • Training algorithm / mode: On-policy distillation (OPD) / orch.
  • Training data: 75% new verified SWE-rebench tasks + 25% historical SWE-bench train350, by prompt groups per update.
  • Local checkpoint: u109. uN is the recorded local checkpoint iteration label, not the original HF OPD step number. Solo350 continues numbering from the earlier local63 checkpoint.
  • Optimizer: weights-only resumes use a fresh optimizer; this run includes restarts.

Teacher: allenai/tmax-27b; MiniMax-M2.7 coordinates training work orders. OPD teacher KL coefficient is 1.0; reference-policy KL is 0. Source pool: 774 accepted new tasks plus 350 historical tasks, sampled 75/25 by prompt groups. Runtime image exclusions are recorded separately in TRAINING_CONFIG.json.

Training setting Value
Learning rate 5e-07
Samples per update 64
Group size 4
Prompt groups per update 16
Reference-policy KL coefficient 0.0
Maximum weight staleness 1
Save interval 4 updates
Worker thinking disabled

Exact parent revisions, dataset hashes, archived training settings and native checkpoint identity are in CHECKPOINT_PROVENANCE.json and TRAINING_CONFIG.json.

Evaluation settings

150 held-out tasks, 32 episodes per job, 2 CPUs and 10 GiB per sandbox; worker thinking disabled. Worker/coordinator temperatures: 0.2; maximum worker turns: 64; response budgets: 8192/4096 tokens; verifier timeout: 2400 seconds. CPU-quota profiles additionally set Django/numeric thread defaults to 2. Exact per-branch settings, task outcomes, fingerprints, repetitions and resource profiles are in EVAL_RESULTS.json.

Download a checkpoint

from huggingface_hub import snapshot_download
path = snapshot_download("CharlieLLL/Qwen3.5-9B-Coding-MixedOPD-TMAX27B", revision="u109")

Use a serving stack with Qwen3.5 support and the included tokenizer/chat template. These are full BF16 weights plus tokenizer/processor assets; optimizer state is not included. Evaluation covers text-based coding tasks; inherited vision configuration does not establish evaluated multimodal behavior. License: Apache-2.0; see LICENSE.

Downloads last month
24
Safetensors
Model size
9B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for CharlieLLL/Qwen3.5-9B-Coding-MixedOPD-TMAX27B

Finetuned
(13)
this model