Qwen3.5-9B-Coding-Raw-SoloRL-Frontier124-u309
Full Qwen3.5-9B checkpoint from Solo GRPO, local checkpoint u309.
训练来源:solo-rl-frontier124-raw-g8-fast-20260916,checkpoint u309;下方成绩为 orch 评测。
Training run and lineage
- Run name:
from-raw_solo-grpo_frontier124_b128g8_lr5e-7_kl1e-3_stale4_u110_sample-refill_tok32k_restart1-rl - W&B: 815ab3ed
- Lineage: Raw Qwen3.5-9B → Solo GRPO on historical frontier124.
- Training mode:
solo. The Solo training phase has no teacher or orchestrator. - Step numbering: u309 is this campaign's local checkpoint label, not original HF OPD step309.
- Data: Historical frontier124 variance/std-selected SWE-bench training subset.
- Optimizer: weights-only resumes start a fresh optimizer; the run includes restarts.
| Training setting | Value |
|---|---|
| Learning rate | 5e-07 |
| Samples per update | 128 |
| Group size | 8 |
| Prompt groups per update | 16 |
| Reference-policy KL coefficient | 0.001 |
| Maximum weight staleness | 4 |
| Save interval | 5 updates |
| Worker thinking | disabled |
Exact starting-model revisions, dataset hashes, and selected archived training settings are in CHECKPOINT_PROVENANCE.json and TRAINING_CONFIG.json.
Evaluation: M2.7 orchestration, held-out eval150
This checkpoint is the small coding worker; MiniMax-M2.7 is the coordinator. This is an orch result, including for checkpoints trained with Solo RL.
| Same-cohort evaluation | Resolved |
|---|---|
| This checkpoint | 87/150 (58.0%) |
| Actual starting-model control | 84/150 (56.0%) |
| Difference | +3 tasks (+2.0 percentage points) |
Both runs completed all 150 tasks and passed the evaluator audit, with zero final infrastructure errors. Model/test failures and verifier timeouts count as unresolved. The protocol uses 32 episodes per job, 2 CPUs and 10 GiB per sandbox, and CPU-quota-aware Django/numeric thread defaults. Worker thinking is disabled. Both model temperatures are 0.2; maximum worker turns 64, worker response budget 8192, coordinator response budget 4096, and verifier timeout 2400 seconds. Exact settings, fingerprints and paired task outcomes are in EVAL_RESULTS.json.
Candidates were selected using historical eval150 results. A single paired score does not establish a reproducible training gain. This is a 150-task subset under our orchestration harness, not an official full SWE-bench leaderboard score. Different historical resource protocols are not pooled into this comparison.
Download and use
from huggingface_hub import snapshot_download
path = snapshot_download("CharlieLLL/Qwen3.5-9B-Coding-Raw-SoloRL-Frontier124-u309")
Use a serving stack with Qwen3.5 support and the included tokenizer/chat template. The reported coding evaluation disables worker thinking. These weights were trained and evaluated for text-based coding tasks; inherited vision configuration is not a claim that multimodal behavior was evaluated. The release contains model weights, tokenizer/processor assets and experiment metadata; optimizer state is not included.
License: Apache-2.0; see LICENSE.
- Downloads last month
- 30