qwen36-twla-adaptive-init3-target2

Public research checkpoint produced by routed-expert TWLA mixed-precision optimization. GPQA was not used for calibration, sensitivity ranking, allocation, stopping, or checkpoint selection.

  • Quantization unit: one routed expert in one MoE layer
  • Number of units: 40 layers x 256 experts = 10,240
  • Objective: validation NLL plus a lambda-weighted average logical-bit term
  • Search mode: proxy_prefix
  • Initial level: 3
  • Target routed-expert bits: 2.0
  • Final routed-expert bits: 1.7982684843261683

Reproducibility metadata is stored in optimization_summary.json and precision_map.json. The exact optimizer and GPQA inference sources used in this workspace are included under code/.

Downloads last month
-
Safetensors
Model size
35B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support