decision-model-large
| model | Decision Index | balanced_raw | head-to-head (43 benchmarks) |
|---|---|---|---|
| decision-model-tiny (0.8B) | 22.60 | 41.76 | 3 W / 1 T / 39 L |
| decision-model-small (4B) | 40.66 | 54.85 | 5 W / 2 T / 36 L |
| decision-model-large (176B-A6B, this repo) | 58.52 | 68.32 | 25 W / 1 T / 17 L |
| decision-model-extra-large (321B, LoRA) | 57.79 | 67.83 | 2 W / 0 T / 36 L |
Wins vs the 321B reference on identical cases: When2Call MCQ 0.785 vs 0.559 (+22.6pp), SimpleBench 0.500 vs 0.200 (+30pp), CLINC150+OOS macro-F1 0.878 vs 0.802 (+7.6pp), GSM8K 0.844 vs 0.836, GPQA Diamond 0.561 vs 0.520, POP909-CL 0.347 vs 0.196 (+15.1pp), VAST macro-F1 0.642 vs 0.529 (+11.3pp), ForecastBench Brier 0.162 vs 0.306 (lower is better). Full 155,390-row Decision Index (edition 0.2.1), 5 errors (0.003%).
Qwen3.8-Flash-Next-FP8 (176B-A6B GDN + QSA hybrid MoE, 3:1 GatedDeltaNet:QSA, 1024-expert MoE top-10, hyper-connections, PLE n-gram embedding) + merged rank-32 LoRA (decision-train-v2 recipe: 534k rows, lr 1.5e-5, 2 epochs, cosine, seed 7) + the jev pointer head (head_dim 128), formatted for the jev_sglang out-of-tree sglang plugin family (Qwen4ExpForConditionalGeneration decision plugin).
Contents
model-XXX-of-131.safetensors— Qwen3.8-Flash-Next-FP8 backbone with the LoRA delta merged (W += (A@B).T · (α/r), fp32 accumulate, bf16 out); routed experts stay fp8 with 128×128 block scales byte-identical to the base checkpointhead.safetensors— the pointer head, exactly 5 fp32 keys:head.q_proj.{weight,bias},head.k_proj.{weight,bias},head.log_scaleconfig.json— arch nameQwen4ExpForConditionalGeneration(the plugin REPLACES that registry entry) +language_model_only: true,jev_head_dim: 128,jev_dec_delim_token_id: 52679,jev_opt_delim_token_id: 17317, fp8 quantization config withhead.inmodules_to_not_convertmodel.safetensors.index.json— includes the head keys (sglang's iterator reads only index-listed files)
Serving
Requires the jev_sglang plugin package (not included) mounted via
SGLANG_EXTERNAL_MODEL_PACKAGE=jev_sglang, launched with
--is-embedding --chunked-prefill-size -1 --tp-size 2 --ep-size 2 --mem-fraction-static 0.72 --max-running-requests 48 --max-prefill-tokens 16384 --max-mamba-cache-size 480 --linear-attn-prefill-backend flashinfer --linear-attn-decode-backend flashinfer --mamba-ssm-dtype bfloat16 --context-length 262144
with PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True. The Qwen4ExpForConditionalGeneration plugin replaces the upstream registry entry, enables language_model_only (text-only decision model), and applies the pointer head
(u = q_proj(hidden[dec]), v = k_proj(hidden[opt_end]),
scores = v·u/√d · exp(log_scale), raw logits — temperature is the adapter's job, T*=1.0 unfitted).
Key constraints (see jev_sglang/qwen4_exp_decision.py):
--chunked-prefill-size -1— the plugin's in-band delimiter scan is value-dependent; chunk tails are unscannable--ep-size 2(matching TP) — routed experts cannot TP-split at fp8 block_n=128 (640 intermediate → 320/GPU, not divisible)- Shared-expert fusion disabled via the plugin's
shared_experts_fusion_disable_reason(ships bf16; fused path forces fp8 block quant) head.inmodules_to_not_convert— the fp8 quant pass eats custom heads unless excluded- Dual-naming in
modules_to_not_convert(bothmodel.language_model.*andmodel.*variants) — sglang renames before quant-dispatch head.*keys inmodel.safetensors.index.json— sglang only iterates index-listed files- PLE n-gram table offloaded to pinned host RAM (
ple_offload_embedding: true)
Evaluation
Full 155,390-row Decision Index (edition 0.2.1): index 58.52 (balanced_raw 68.32, breadth 57.68) vs 57.79 for a 321B MoE reference stack on identical cases (25 W / 1 T / 17 L). Area breakdown: Knowledge 0.482, Language 0.637, Retrieval 0.584, Tools 0.754, Arts 0.411.
- Downloads last month
- -