MN9-SHORT-515K-RR3C

LoRA adapter (MN9-SHORT-515K) from the DDC-v3.1 MN9 RR3 short-pool 4-arm experiment. Winner arm: MN9-A-RR3C-SHORT @ ~515K active tokens.

This repo ships the adapter only (~505 MB). Load with base model Qwen/Qwen3-4B-Base. Canonical R2 release: MN9-SHORT-515K-RR3-4ARM-v1 (full 4-arm archive). Do not confuse with a full merged model.

Highlights

Item Value
Codename MN9-SHORT-515K
Winner arm MN9-A-RR3C-SHORT
Base Qwen/Qwen3-4B-Base @ 906bfd4 (pure base, no warm-start)
Training mode NoThink / short visible CoT
Eval AIME24+AIME25, seeds 42–45, EvalScope reviews
Score 36/240 (acc 0.15); BaseFalse 21/240 (+15)
LoRA r=64, α=128, dropout=0; targets q/k/v/o/gate/up/down
Realized active tokens 534970 @ update-16
Adapter SHA256 683a178fd516323c92ade9b1c3e4f8bc3551f774ceabafbc83e7f138052f49b1

Per-seed (winner): 42:9/60, 43:8/60, 44:8/60, 45:11/60.

Four-arm ranking (correct/240): A 36 · B 33 · C 33 · D 32.

Quick load

from peft import PeftModel
from transformers import AutoModelForCausalLM, AutoTokenizer

base_id = "Qwen/Qwen3-4B-Base"
adapter_id = "modrill/MN9-SHORT-515K"

tok = AutoTokenizer.from_pretrained(base_id, trust_remote_code=True)
model = AutoModelForCausalLM.from_pretrained(
    base_id, torch_dtype="auto", device_map="auto", trust_remote_code=True
)
model = PeftModel.from_pretrained(model, adapter_id)
model.eval()

Or see load_example.py in this repo.

Training contract (summary)

  • Length filter: loss_active_tokens < 2048 (no ≥2048 fill / trunc / summarize)
  • Mix target S1/S2/S0 = 0.75/0.15/0.10 (realized ≈ 0.750/0.150/0.100; 441 rows)
  • Runner: DDC_TOKEN_NORMALIZED_RUNNER_V1
  • LR 2e-5 constant, 0 warmup; ~32K active tokens/update; 16 optimizer updates
  • Eval mode: nothink (enable_thinking=false)

Data / protocol

Training data and render protocol live in the DDC-v3.1 experiment tree (runs/mn9_rr3_short_4arm_20260810, render fixture under runs/data_pipeline/MN9/). This Hub package is the inference-usable winner adapter + card, not the full dataset dump.

License

Follow the base model license: Apache 2.0 as used by Qwen/Qwen3-4B-Base. Adapter weights are released under the same terms for research/reproduction of DDC MN9 results.

Provenance

  • Experiment: MN9-RR3-SHORT-4ARM-515K
  • Local canon root: DDC-v3.1/runs/mn9_rr3_short_4arm_20260810
  • Winner symlink: winner/MN9-SHORT-515K → arm A milestone 515K adapter
  • R2: r2mtx:mxx/DDC-v3.1/releases/MN9-SHORT-515K-RR3-4ARM-v1/
  • Registry status: BASELINE (WINNER_REGISTRY)
Downloads last month
11
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for modrill/MN9-SHORT-515K

Adapter
(73)
this model