Qwen3.5-9B-Coding-Raw-SoloRL-Frontier124-u309

Full Qwen3.5-9B checkpoint from Solo GRPO, local checkpoint u309. 训练来源:solo-rl-frontier124-raw-g8-fast-20260916,checkpoint u309;下方成绩为 orch 评测。

Training run and lineage

  • Run name: from-raw_solo-grpo_frontier124_b128g8_lr5e-7_kl1e-3_stale4_u110_sample-refill_tok32k_restart1-rl
  • W&B: 815ab3ed
  • Lineage: Raw Qwen3.5-9B → Solo GRPO on historical frontier124.
  • Training mode: solo. The Solo training phase has no teacher or orchestrator.
  • Step numbering: u309 is this campaign's local checkpoint label, not original HF OPD step309.
  • Data: Historical frontier124 variance/std-selected SWE-bench training subset.
  • Optimizer: weights-only resumes start a fresh optimizer; the run includes restarts.
Training setting Value
Learning rate 5e-07
Samples per update 128
Group size 8
Prompt groups per update 16
Reference-policy KL coefficient 0.001
Maximum weight staleness 4
Save interval 5 updates
Worker thinking disabled

Exact starting-model revisions, dataset hashes, and selected archived training settings are in CHECKPOINT_PROVENANCE.json and TRAINING_CONFIG.json.

Evaluation: M2.7 orchestration, held-out eval150

This checkpoint is the small coding worker; MiniMax-M2.7 is the coordinator. This is an orch result, including for checkpoints trained with Solo RL.

Same-cohort evaluation Resolved
This checkpoint 87/150 (58.0%)
Actual starting-model control 84/150 (56.0%)
Difference +3 tasks (+2.0 percentage points)

Both runs completed all 150 tasks and passed the evaluator audit, with zero final infrastructure errors. Model/test failures and verifier timeouts count as unresolved. The protocol uses 32 episodes per job, 2 CPUs and 10 GiB per sandbox, and CPU-quota-aware Django/numeric thread defaults. Worker thinking is disabled. Both model temperatures are 0.2; maximum worker turns 64, worker response budget 8192, coordinator response budget 4096, and verifier timeout 2400 seconds. Exact settings, fingerprints and paired task outcomes are in EVAL_RESULTS.json.

Candidates were selected using historical eval150 results. A single paired score does not establish a reproducible training gain. This is a 150-task subset under our orchestration harness, not an official full SWE-bench leaderboard score. Different historical resource protocols are not pooled into this comparison.

Download and use

from huggingface_hub import snapshot_download
path = snapshot_download("CharlieLLL/Qwen3.5-9B-Coding-Raw-SoloRL-Frontier124-u309")

Use a serving stack with Qwen3.5 support and the included tokenizer/chat template. The reported coding evaluation disables worker thinking. These weights were trained and evaluated for text-based coding tasks; inherited vision configuration is not a claim that multimodal behavior was evaluated. The release contains model weights, tokenizer/processor assets and experiment metadata; optimizer state is not included.

License: Apache-2.0; see LICENSE.

Downloads last month
30
Safetensors
Model size
9B params
Tensor type
F32
·
BF16
·
Video Preview
loading

Model tree for CharlieLLL/Qwen3.5-9B-Coding-Raw-SoloRL-Frontier124-u309

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(996)
this model