Vishva007/Qwen3.5-9B-W4A16-AutoRound-LLM-Compressor

This is a W4A16 (4-bit weight, 16-bit activation) LLM-Compressor-format quantized version of Qwen/Qwen3.5-9B, produced using AutoRound — Intel's sign gradient descent based quantization method designed for production-grade accuracy retention.

  • Vision Tower (quant_nontext_module): False (Kept in BF16 to preserve visual reasoning and OCR precision)
  • Special Modules (layer_config): Multi-Token Prediction (mtp, mtp.fc) kept in native bfloat16

Quantization Details

Parameter Value
Method AutoRound (W4A16, LLM-Compressor format)
Group Size 32
Symmetric Yes
Iterations 800
Calibration Samples 512
Sequence Length 4096
Torch Compile Enabled

Key Notes

  • LLM-Compressor format — Exported in the standard LLM-Compressor format for broad ecosystem compatibility.
  • High accuracy configuration — 800 iterations with 512 calibration samples targets production-grade quality with minimal degradation from the base model.
  • W4A16 — Weights are quantized to 4-bit integers; activations remain in FP16 for inference stability.
  • ~50% memory reduction compared to the FP16 base model, enabling deployment on consumer and mid-range GPUs.

Usage

This model is compatible with transformers, LLM-Compressor, vLLM, and SGLang — any backend supporting LLM-Compressor-format weights works out of the box. For full model details, architecture, and capabilities, refer to the base model page.

Downloads last month
-
Safetensors
Model size
4B params
Tensor type
I32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Vishva007/Qwen3.5-9B-W4A16-AutoRound-LLM-Compressor

Finetuned
Qwen/Qwen3.5-9B
Quantized
(484)
this model

Collection including Vishva007/Qwen3.5-9B-W4A16-AutoRound-LLM-Compressor