Vishva007/Qwen3.5-0.8B-W4A16-AutoRound-LLM-Compressor
This is a W4A16 (4-bit weight, 16-bit activation) GPTQ-format quantized version of Qwen/Qwen3.5-0.8B, produced using AutoRound — Intel's sign gradient descent based quantization method designed for production-grade accuracy retention.
Quantization Details
| Parameter |
Value |
| Method |
AutoRound (W4A16, GPTQ format) |
| Group Size |
128 |
| Symmetric |
Yes |
| Iterations |
1000 |
| Calibration Samples |
512 |
| Sequence Length |
2048 |
| Torch Compile |
Enabled |
Key Notes
- GPTQ format — Exported in the standard GPTQ format for broad ecosystem compatibility.
- Ultra-high accuracy configuration — 1000 iterations with 512 calibration samples ensures near-lossless quantization, especially critical at this model scale where parameter budget is tight.
- W4A16 — Weights are quantized to 4-bit integers; activations remain in FP16 for inference stability.
- Extremely lightweight — The quantized 0.8B model is suitable for edge deployment, low-latency inference, and resource-constrained environments.
- ~50% memory reduction compared to the FP16 base model.
Usage
This model is compatible with transformers, AutoGPTQ, vLLM, and SGLang — any backend supporting GPTQ-format weights works out of the box. For full model details, architecture, and capabilities, refer to the base model page.
🚀 Deploy on RunPod
One-click launch environments pre-configured with PyTorch, CUDA, and dependencies for fine-tuning or quantization.
🎁 Need GPU compute? Sign up via RunPod and get $5–$500 in free credits when you add your first $10.
PyTorch 2.14
| Template |
CUDA Version |
Docker Image |
Template ID |
Deploy |
| PyTorch 2.14 (CUDA 12.6) |
12.6 |
vishva123/cuda-12.6-pytorch-2.14-runpod |
d7lxsa4w9m |
 |
| PyTorch 2.14 (CUDA 13.0) |
13.0 |
vishva123/cuda-13.0-pytorch-2.14-runpod |
yk0y6j6rpg |
 |
| PyTorch 2.14 (CUDA 13.2) |
13.2 |
vishva123/cuda-13.2-pytorch-2.14-runpod |
gsp4gwx0nw |
 |
PyTorch 2.13
| Template |
CUDA Version |
Docker Image |
Template ID |
Deploy |
| PyTorch 2.13 (CUDA 12.6) |
12.6 |
vishva123/cuda-12.6-pytorch-2.13-runpod |
gmlupxnxfk |
 |
| PyTorch 2.13 (CUDA 13.0) |
13.0 |
vishva123/cuda-13.0-pytorch-2.13-runpod |
y3j8xvk4f4 |
 |
| PyTorch 2.13 (CUDA 13.2) |
13.2 |
vishva123/cuda-13.2-pytorch-2.13-runpod |
vigpissn5w |
 |
PyTorch 2.12
| Template |
CUDA Version |
Docker Image |
Template ID |
Deploy |
| PyTorch 2.12 (CUDA 12.6) |
12.6 |
vishva123/cuda-12.6-pytorch-2.12-runpod |
ctmz86zmf0 |
 |
| PyTorch 2.12 (CUDA 13.0) |
13.0 |
vishva123/cuda-13.0-pytorch-2.12-runpod |
qjko5yiwzi |
 |
| PyTorch 2.12 (CUDA 13.2) |
13.2 |
vishva123/cuda-13.2-pytorch-2.12-runpod |
ifg6xmye0f |
 |