RedHatAI/Qwen3.8-Flash-Next-NVFP4

This is a quantized version of Qwen/Qwen3.8-Flash-Next with MoE layers quantized to NVFP4. The model was calibrated using 1024 samples from perfectblend.

Usage

This model is intended for deployment with vLLM. You can serve the model using

vllm serve RedHatAI/Qwen3.8-Flash-Next-NVFP4 \
  --tensor-parallel-size 4 \
  --enable-expert-parallel \

Evaluation

inspect eval hf/Idavidrein/gpqa/diamond \
  --model vllm/RedHatAI/Qwen3.8-Flash-Next-NVFP4 \
  --reasoning-effort xhigh \
  --model-base-url http://localhost:8000/v1 \
  -M client_timeout=2400 \
  --token-limit 100000 \
  --retry-on-error=2
Benchmark Qwen/Qwen3.8-Flash-Next RedHatAI/Qwen3.8-Flash-Next-NVFP4
GPQA Diamond 91.7 90.9
Downloads last month
138
Safetensors
Model size
180B params
Tensor type
BF16
·
U8
·
I64
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RedHatAI/Qwen3.8-Flash-Next-NVFP4

Quantized
(236)
this model