Qwen3-4B-Computer-Science-AWQ
Overview
Qwen3-4B-Computer-Science-AWQ is the AWQ-calibrated quantized release of Qwen3-4B-Computer-Science.
Weights are stored using the Compressed-Tensors format with 4-bit asymmetric group-wise quantization (W4A16). The checkpoint was produced using Activation-aware Weight Quantization (AWQ) calibration and validated by successful quantization, checksum verification, and CPU inference.
Model Information
| Property | Value |
|---|---|
| Base Model | Qwen/Qwen3-4B |
| Model Family | Qwen3-4B-Computer-Science |
| Quantization | AWQ |
| Storage Format | Compressed-Tensors |
| Weight Precision | INT4 |
| Activation Precision | FP16 / BF16 |
| Quantization Scheme | W4A16 |
| Group Size | 128 |
| Weight Quantization | Asymmetric |
| lm_head | Excluded from Quantization |
| Language | English |
| License | Apache-2.0 |
Training Data
The base model was instruction tuned using permissively licensed datasets.
| Dataset | Configuration | License |
|---|---|---|
| HuggingFaceTB/smoltalk | smol-magpie-ultra | Apache-2.0 |
| agentica-org/DeepCoder-Preview-Dataset | primeintellect | MIT |
Dataset Size
| Split | Samples |
|---|---|
| Training | 60,989 |
| Evaluation | 512 |
Intended Use
This model is intended for:
- Software engineering
- Programming
- Code generation
- Debugging
- Code review
- Algorithm implementation
- Computer science education
- General technical reasoning
Quantization
This release was generated using Activation-aware Weight Quantization (AWQ).
The resulting checkpoint stores weights using packed 4-bit group-wise asymmetric quantization.
| Parameter | Value |
|---|---|
| Weight Format | Packed INT4 |
| Group Size | 128 |
| Symmetric | No |
| Observer | memoryless_minmax |
| Compression Format | Compressed-Tensors |
Runtime Compatibility
This checkpoint uses the Compressed-Tensors format.
It is intended for runtimes that support Compressed-Tensors models.
Validation performed for this release:
- Successful AWQ calibration
- Successful model serialization
- CPU inference
- SHA256 verification of release artifacts
Loading this checkpoint with standard Transformers may decompress weights during execution depending on the runtime and available hardware.
Usage
Transformers
from transformers import AutoModelForCausalLM, AutoTokenizer
model = AutoModelForCausalLM.from_pretrained(
"Irfanuruchi/Qwen3-4B-Computer-Science-AWQ",
device_map="auto",
dtype="auto",
)
tokenizer = AutoTokenizer.from_pretrained(
"Irfanuruchi/Qwen3-4B-Computer-Science-AWQ"
)
Release Artifacts
model.safetensors
config.json
generation_config.json
recipe.yaml
tokenizer.json
tokenizer_config.json
chat_template.jinja
SHA256SUMS
LICENSE
README.md
Integrity Verification
Every release artifact includes a SHA256 checksum.
Verify downloaded files:
sha256sum -c SHA256SUMS
Validation
The published checkpoint was verified before release.
Completed validation:
- AWQ calibration completed successfully
- Quantized checkpoint generated successfully
- CPU inference completed successfully
- SHA256 checksums verified
Limitations
- Quantization may affect output quality compared to the BF16 checkpoint.
- Runtime support depends on the inference engine.
- GPU memory requirements depend on whether the runtime executes directly on compressed weights or decompresses them during inference.
License
Base model:
- Apache-2.0
Training datasets:
- Apache-2.0
- MIT
This repository is distributed under the Apache-2.0 License.
Acknowledgements
- Alibaba Qwen Team
- Hugging Face
- vLLM Project
- LLM Compressor Project
- SmolTalk Contributors
- DeepCoder Contributors
Citation
@software{uruci2026qwen3csawq,
title={Qwen3-4B-Computer-Science-AWQ},
author={Irfan Uru莽i},
year={2026},
publisher={Hugging Face}
}
- Downloads last month
- 35