You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Qwen3.8 Flash-Next NVFP4 with calibrated FP8 KV (Spark candidate)

This is a manually gated evaluation release of the NVIDIA-derived nvidia/Qwen3.8-Flash-Next-NVFP4 checkpoint. Every access request must be reviewed by the repo owner. The repository is not a public production promotion.

What is in this release

  • Native NVIDIA/ModelOpt mixed-precision weights: routed experts use NVFP4 (group size 16), the MTP expert path uses FP8, and the PLE n-gram embedding uses FP8. Protected model regions retain their source precision.
  • Calibrated FP8 KV metadata: 26 FP32 scales covering 13 main/MTP QSA owners.
  • vLLM serving recipe tested on DGX Spark GB10: eager execution, MTP3, prefix caching, 262,144-token maximum length, and an 8,192-token output reserve.
  • CPU-deduplicated PLE reader selected for the qualified private deployment. GPU-deduplication was evaluated and found to have no benefits.

Evidence and limits

The candidate loaded and served privately on Spark-01. Against the original reader (Tony) , the same weights and runtime measured approximately 14% lower mixed-C4 p95 latency. The authored campaign screens were retained (50/60 coding cases, 9/12 task cases, 10/12 matrix cases, and 3/3 image cases), and the campaign exercised the 262,144-token envelope.

This release does not claim a served BF16-teacher KL gate. The mapped GPU PLE path was implemented and measured, but its bounded serving comparison did not show a reliable end-to-end speed gain over CPU dedup.

Use the exact vLLM/runtime compatibility documented. Do not infer support for other GPUs or serving stacks from this artifact.

Provenance

  • NVIDIA source revision: fc694b54fb0174e0913e6adf86691ef85a4ead47
  • vLLM source pin: 8a728663c1c3eeace834a95f5654fa653cc1998c
  • Candidate manifest SHA-256: fcbce8545cff454027e98450a60ce0f139e8cad19ac4f127d2a085ce669ecaf2
  • Checkpoint manifest SHA-256: 981d8b768a08a0a823243f149d57991bf702ccf4a8b2e876c899c9001a22fb75
  • Calibration receipt SHA-256: 3bc5c7d188b87f86ae1011b6109786bc82580da926088fcff7e640073f3044fd

The full source, receipts, calibration-source audit, and human reports remain in the accompanying Windows/NUC release bundle.

The upstream NVIDIA model card declares nvidia-open-model-license; the underlying Qwen model declares qwen-community-1.0. Review and comply with both upstream licenses before requesting access.

Downloads last month
-
Safetensors
Model size
120B params
Tensor type
BF16
·
U8
·
I64
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for JasonW2025/Qwen3.8-Flash-Next-NVFP4-Calibrated-KV

Quantized
(3)
this model