decision-model-large

model Decision Index balanced_raw head-to-head (43 benchmarks)
decision-model-tiny (0.8B) 22.60 41.76 3 W / 1 T / 39 L
decision-model-small (4B) 40.66 54.85 5 W / 2 T / 36 L
decision-model-large (176B-A6B, this repo) 58.52 68.32 25 W / 1 T / 17 L
decision-model-extra-large (321B, LoRA) 57.79 67.83 2 W / 0 T / 36 L

Wins vs the 321B reference on identical cases: When2Call MCQ 0.785 vs 0.559 (+22.6pp), SimpleBench 0.500 vs 0.200 (+30pp), CLINC150+OOS macro-F1 0.878 vs 0.802 (+7.6pp), GSM8K 0.844 vs 0.836, GPQA Diamond 0.561 vs 0.520, POP909-CL 0.347 vs 0.196 (+15.1pp), VAST macro-F1 0.642 vs 0.529 (+11.3pp), ForecastBench Brier 0.162 vs 0.306 (lower is better). Full 155,390-row Decision Index (edition 0.2.1), 5 errors (0.003%).

Qwen3.8-Flash-Next-FP8 (176B-A6B GDN + QSA hybrid MoE, 3:1 GatedDeltaNet:QSA, 1024-expert MoE top-10, hyper-connections, PLE n-gram embedding) + merged rank-32 LoRA (decision-train-v2 recipe: 534k rows, lr 1.5e-5, 2 epochs, cosine, seed 7) + the jev pointer head (head_dim 128), formatted for the jev_sglang out-of-tree sglang plugin family (Qwen4ExpForConditionalGeneration decision plugin).

Contents

  • model-XXX-of-131.safetensors — Qwen3.8-Flash-Next-FP8 backbone with the LoRA delta merged (W += (A@B).T · (α/r), fp32 accumulate, bf16 out); routed experts stay fp8 with 128×128 block scales byte-identical to the base checkpoint
  • head.safetensors — the pointer head, exactly 5 fp32 keys: head.q_proj.{weight,bias}, head.k_proj.{weight,bias}, head.log_scale
  • config.json — arch name Qwen4ExpForConditionalGeneration (the plugin REPLACES that registry entry) + language_model_only: true, jev_head_dim: 128, jev_dec_delim_token_id: 52679, jev_opt_delim_token_id: 17317, fp8 quantization config with head. in modules_to_not_convert
  • model.safetensors.index.json — includes the head keys (sglang's iterator reads only index-listed files)

Serving

Requires the jev_sglang plugin package (not included) mounted via SGLANG_EXTERNAL_MODEL_PACKAGE=jev_sglang, launched with --is-embedding --chunked-prefill-size -1 --tp-size 2 --ep-size 2 --mem-fraction-static 0.72 --max-running-requests 48 --max-prefill-tokens 16384 --max-mamba-cache-size 480 --linear-attn-prefill-backend flashinfer --linear-attn-decode-backend flashinfer --mamba-ssm-dtype bfloat16 --context-length 262144 with PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True. The Qwen4ExpForConditionalGeneration plugin replaces the upstream registry entry, enables language_model_only (text-only decision model), and applies the pointer head (u = q_proj(hidden[dec]), v = k_proj(hidden[opt_end]), scores = v·u/√d · exp(log_scale), raw logits — temperature is the adapter's job, T*=1.0 unfitted).

Key constraints (see jev_sglang/qwen4_exp_decision.py):

  • --chunked-prefill-size -1 — the plugin's in-band delimiter scan is value-dependent; chunk tails are unscannable
  • --ep-size 2 (matching TP) — routed experts cannot TP-split at fp8 block_n=128 (640 intermediate → 320/GPU, not divisible)
  • Shared-expert fusion disabled via the plugin's shared_experts_fusion_disable_reason (ships bf16; fused path forces fp8 block quant)
  • head. in modules_to_not_convert — the fp8 quant pass eats custom heads unless excluded
  • Dual-naming in modules_to_not_convert (both model.language_model.* and model.* variants) — sglang renames before quant-dispatch
  • head.* keys in model.safetensors.index.json — sglang only iterates index-listed files
  • PLE n-gram table offloaded to pinned host RAM (ple_offload_embedding: true)

Evaluation

Full 155,390-row Decision Index (edition 0.2.1): index 58.52 (balanced_raw 68.32, breadth 57.68) vs 57.79 for a 321B MoE reference stack on identical cases (25 W / 1 T / 17 L). Area breakdown: Knowledge 0.482, Language 0.637, Retrieval 0.584, Tools 0.754, Arts 0.411.

Downloads last month
-
Safetensors
Model size
177B params
Tensor type
BF16
·
F8_E4M3
·
I64
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for someguystrainingmodels/decision-model-large

Quantized
(11)
this model