RedHatAI/Qwen3.8-2.4T-A95B-NVFP4-FP8

This is a quantized version of Qwen/Qwen3.8-2.4T-A95B with MoE layers quantized to NVFP4 and attention layers quantized to FP8 block

Usage

This model is intended for deployment with vLLM. You can serve the model using

vllm serve RedHatAI/Qwen3.8-2.4T-A95B-NVFP4 \
    --data-parallel-size 8 \
    --enable-expert-parallel 8 \
    --reasoning-parser qwen3 \
    --max-num-seqs 140

NOTE: Because of the need for data parallelism, this model may use more model memory for replicated attention layers. This can lead to less memory for kv cache and cache preemption when handling large concurrency. For high concurrency tasks, consider using RedHatAI/Qwen3.8-2.4T-A95B-NVFP4.

Creation Process

This model was quantized using LLM Compressor, see https://github.com/vllm-project/llm-compressor/blob/main/docs/key-models/qwen3.5/nvfp4-moe-example.md

Evaluation

inspect eval hf/Idavidrein/gpqa/diamond
  --model vllm/RedHatAI/Qwen3.8-2.4T-A95B-NVFP4-FP8
  --reasoning-effort xhigh
  --model-base-url http://localhost:8000/v1
  -M client_timeout=2400
  --token-limit 100000
  --retry-on-error=2
Benchmark Qwen/Qwen3.8-2.4T-A95B RedHatAI/Qwen3.8-2.4T-A95B-NVFP4-FP8 RedHatAI/Qwen3.8-2.4T-A95B-NVFP4 Inferact/Qwen3.8-2.4T-A95B-NVFP4
GPQA DIamond 92.6 93.1 92.9 92.9
Downloads last month
54
Safetensors
Model size
1.4T params
Tensor type
BF16
·
U8
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RedHatAI/Qwen3.8-2.4T-A95B-NVFP4-FP8

Quantized
(25)
this model