Qwen3.8-Flash-Next NVFP4 with NVFP4 PLE

This repository is a serving derivative of Inferact/Qwen3.8-Flash-Next-NVFP4. It keeps the ModelOpt NVFP4 backbone unchanged and converts only the BF16 position-learning enhancement (PLE) n-gram table to NVFP4.

This repository contains model weights and Hugging Face metadata only. It does not include runtime code, shell scripts, Docker files, or result artifacts.

The packed PLE format requires an external qwen38-nvfp4-ple runtime plugin. Stock vLLM and Transformers cannot load it. The checkpoint retains the original vision tower.

Storage

Artifact Exact bytes GiB
Source checkpoint payload 182,738,792,440 170.189
Source BF16 PLE table 102,400,491,520 95.368
Converted NVFP4 PLE table 28,800,138,752 26.822
Complete converted checkpoint 109,138,439,672 101.643

The conversion saves 73,600,352,768 bytes, or 68.546 GiB. It does not use REAP or remove experts.

Why this fits on one DGX Spark

This release targets one NVIDIA DGX Spark because its GB10 CPU and GPU share 128 GB of LPDDR5x unified system memory. NVIDIA documents the hardware and memory layout. The GPU executes the NVFP4 backbone. The plugin keeps the 26.822 GiB packed PLE table CPU-offloaded and decodes only requested rows to BF16. Both allocations draw from the same physical memory pool.

In a measured full multimodal TP1 run, vLLM reported 73.97 GiB during model loading and reserved a 2 GiB KV cache. The vision tower added 0.85 GiB over text-only serving. The PLE worker has been tested only with vLLM's mp executor at TP1.

API startup took about 639 seconds. After image and text requests, the host had 19.8 GiB available with 14.0 GiB of system-wide swap in use. The default 4,096-token context, single sequence, and 2 GiB KV cache are conservative. This demonstrates fit; it is not a throughput claim.

External runtime contract

Component Required or tested value
vLLM image vllm/vllm-openai:qwen38-flash-next@sha256:fc120ece0a388cc0aa1caad4a9f1cd92113484ab7ec2fd0efadd62585be05bf8
vLLM version 0.1.dev20073+g8e685d198
PLE mode CPU offload with BF16 lookup output
Serving model Full Qwen4ExpForConditionalGeneration wrapper
Multimodal limits One image, no video, 65,536-to-262,144-pixel budget
Executor mp, including at TP1
Validated GPU NVIDIA GB10, compute capability 12.1

The pinned vLLM uni executor does not spawn the PLE CPU worker. The external plugin must select mp and register itself inside vLLM's special PLE subprocess.

PLE format

Format qwen38-nvfp4-ple-v1 stores each group of 16 values as packed E2M1 values, one E4M3 block scale, and one FP32 global decode scale per source PLE shard. The runtime gathers only requested rows and decodes those rows to BF16. It never expands the complete table.

Evaluation

A controlled teacher-forced study used 32 calibration prompts and eight held- out prompts, covering 236 next-token positions. The executable reference was the official FP8 backbone with normal PLE behavior.

These measurements isolate changes to the PLE representation. They do not measure the separate quality difference between the Inferact NVFP4 backbone and the original model. This uploaded checkpoint uses the first row below; the REAP rows are comparison experiments only.

Perturbation Top-1 agreement Reference top-1 in candidate top-5 KL, nats NLL delta, nats Logit cosine
NVFP4 PLE only 91.525% 100.000% 0.015591 +0.019442 0.995399
REAP only, 25% experts removed 87.712% 100.000% 0.042569 +0.043668 0.991027
REAP and NVFP4 PLE 91.102% 99.576% 0.040298 +0.040545 0.990824

The converted checkpoint passed text-only and full multimodal vLLM smoke tests. The multimodal run loaded the vision tower and served PNG inputs. With a neutral prompt, it correctly identified a generated Moon scene, Earth, casual clothing, and the absence of a spacesuit.

Limits

  • The alignment study is not a task-quality benchmark.
  • The runtime checks are smoke tests, not throughput or vision benchmarks.
  • Long-context, concurrent-request, tensor-parallel, and expert-parallel paths have not been validated.
  • The validated image path accepts one image within the default pixel budget. Video is disabled by default and has not been validated.
  • The source MTP weights remain in the checkpoint, but the validated runtime did not enable MTP.
  • The tested host had system-wide swap in use after inference.
  • The PLE decoder uses ordinary PyTorch CPU operations. A fused decoder could reduce lookup latency.

Provenance and license

The backbone comes from Inferact/Qwen3.8-Flash-Next-NVFP4, which derives from Qwen/Qwen3.8-Flash-Next. Only the PLE representation changed here.

The model weights remain under the Qwen Community License 1.0 in LICENSE.

Downloads last month
-
Safetensors
Model size
93B params
Tensor type
I64
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for provsalt/Qwen3.8-Flash-Next-NVFP4-PLE-NVFP4

Quantized
(1)
this model