RedHatAI/Qwen3.8-2.4T-A95B-NVFP4

This is a quantized version of Qwen/Qwen3.8-2.4T-A95B with MoE layers quantized to NVFP4

Usage

This model is intended for deployment with vLLM. You can serve the model using

vllm serve RedHatAI/Qwen3.8-2.4T-A95B-NVFP4 \
    --tensor-parallel-size 8 \
    --enable-expert-parallel 8 \
    --reasoning-parser qwen3

Creation Process

This model was quantized using LLM Compressor, see https://github.com/vllm-project/llm-compressor/blob/main/docs/key-models/qwen3.5/nvfp4-moe-example.md

Downloads last month
53
Safetensors
Model size
1.4T params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for RedHatAI/Qwen3.8-2.4T-A95B-NVFP4

Quantized
(22)
this model