YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Qwen3.8-Flash-Next bf16 reference logits

Ground truth for Splash quality checks, from dev/benchmarks/qwen4exp/reference_logits.py (full-precision checkpoint, float32 on CPU). Each <name>.reference.bin is bf16 logits [tokens x 248320] for <name>.passage.json.

  • code: 1,495-token code prompt; short: 65 tokens; long4k: first 4,096 tokens of the long prompt (crosses the 2,048-token sparse-indexer budget).
  • code.plan.bin: the reference run with the current package's quantization simulated (--quant experts:4,router:8,shared:4,attn:8,gdn_in:8,gdn_ab:8, gdn_out:8,ple:4,head:8,embed:8): 90.2% same pick, KL 0.122.

Score an engine dump: SPLASH_DUMP_PREFILL_LOGITS=/tmp/x.bin build/engine-tests/generate-sample
build/splash.metallib ~/models/qwen38-flash-next-splash 1 "" python dev/benchmarks/qwen4exp/compare_logits.py --passages code.passage.json
--name code --reference code.reference.bin --engine /tmp/x.bin [--begin N --end M] (Use ~/vllm-mlx-env-3.12/bin/python for the reference and compare scripts.)

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support