Qwen-3.5 Collection
Collection
Quantized Qwen3.5 models for efficient image-text understanding (AutoRound W4A16). • 13 items • Updated • 1
This is a W4A16 (4-bit weight, 16-bit activation) LLM-Compressor-format quantized version of Qwen/Qwen3.5-9B, produced using AutoRound — Intel's sign gradient descent based quantization method designed for production-grade accuracy retention.
quant_nontext_module): False (Kept in BF16 to preserve visual reasoning and OCR precision)layer_config): Multi-Token Prediction (mtp, mtp.fc) kept in native bfloat16| Parameter | Value |
|---|---|
| Method | AutoRound (W4A16, LLM-Compressor format) |
| Group Size | 32 |
| Symmetric | Yes |
| Iterations | 800 |
| Calibration Samples | 512 |
| Sequence Length | 4096 |
| Torch Compile | Enabled |
This model is compatible with transformers, LLM-Compressor, vLLM, and SGLang — any backend supporting LLM-Compressor-format weights works out of the box. For full model details, architecture, and capabilities, refer to the base model page.