Qwen3.8-27B-NVFP4

NVFP4 (4-bit, compressed-tensors nvfp4-pack-quantized) weight quantization of Qwen/Qwen3.8-27B for fast inference on NVIDIA Blackwell (sm_120) with vLLM.

Details

  • Base model: Qwen/Qwen3.8-27B. Dense 27B, native vision-language, Gated-DeltaNet hybrid with MTP.
  • Method: weight-only NVFP4, group size 16 (compressed-tensors).
  • Target runtime: vLLM on Blackwell-class GPUs.
  • Relation to base: quantization only. No weights were trained or fine-tuned.

Usage

vllm serve Preyazz/Qwen3.8-27B-NVFP4 --trust-remote-code

Provenance and license

This repository redistributes a quantized copy of Qwen/Qwen3.8-27B. All model capabilities, credit, and the governing license belong to the Qwen team, and the base model's license applies to this quantized derivative. See the original model card for full model documentation, intended use, and limitations.

Provided as-is, without warranty.

Downloads last month
393
Safetensors
Model size
27B params
Tensor type
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Preyazz/Qwen3.8-27B-NVFP4

Base model

Qwen/Qwen3.8-27B
Quantized
(599)
this model