Qwen3-4B Quanto INT8
INT8 weight-quantized version of Qwen/Qwen3-4B using persistent Optimum-Quanto qint8 weights with BF16 activations.
Important: This is a persistent Optimum-Quanto checkpoint. Load it with
optimum.quanto.QuantizedModelForCausalLMas shown below. Do not assume that genericAutoModelForCausalLM, vLLM, SGLang, AWQ, GPTQ, GGUF, or TorchAO loading paths are compatible unless that application explicitly supports persistent Optimum-Quanto checkpoints.
Quantization
- Base model:
Qwen/Qwen3-4B - Persistent Optimum-Quanto
qint8 - BF16 activations
- All Linear layers quantized, including
lm_head - Repository size: approximately 4.484 GiB
- No custom
trust_remote_code=Trueloader required - PAD:
151643 - EOS:
[151645, 151643] - BOS:
151643
Requirements
Validated with:
- Python 3.12.10
- PyTorch 2.12.0+cu130
- Transformers 5.17.0
- Optimum-Quanto
- NVIDIA RTX 4060 Ti 16 GB
Install the model-side dependencies:
pip install "transformers>=5.17.0" optimum-quanto safetensors huggingface_hub
Install PyTorch separately using the build appropriate for your CUDA/CPU environment.
TorchAO warning
TorchAO is not required for this model.
If torchao is installed in the same environment, newer PyTorch versions may
print deprecation messages mentioning KernelPreference,
ScaleCalculationMode, or register_constant(). These originate from the
TorchAO/PyTorch dependency combination, not this Qwen3-4B INT8 checkpoint,
and do not indicate failed model loading or broken quantization.
If another application requires TorchAO, the messages can be ignored. If TorchAO is not needed, removing it eliminated the messages in the tested environment.
Correct loading method
import torch
from transformers import AutoTokenizer
from optimum.quanto import QuantizedModelForCausalLM
model_id = "Redtash1/Qwen3-4B-INT8"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = QuantizedModelForCausalLM.from_pretrained(
model_id,
dtype=torch.bfloat16,
)
# Explicit transfer is recommended for this persistent Quanto checkpoint.
model = model.to("cuda")
model.eval()
During validation, relying on device_map="cuda" alone did not reliably move
the persistent Quanto wrapper to the GPU, so explicit .to("cuda") is
recommended.
Example generation
messages = [
{
"role": "user",
"content": "Write a short description of a quiet mountain lake."
}
]
text = tokenizer.apply_chat_template(
messages,
tokenize=False,
add_generation_prompt=True,
enable_thinking=False,
)
inputs = tokenizer(text, return_tensors="pt").to("cuda")
with torch.inference_mode():
output = model.generate(
**inputs,
max_new_tokens=256,
do_sample=True,
temperature=0.7,
top_p=0.8,
top_k=20,
)
new_tokens = output[0, inputs["input_ids"].shape[1]:]
print(tokenizer.decode(new_tokens, skip_special_tokens=True))
Compatibility
This checkpoint is Optimum-Quanto, not:
- TorchAO
- AWQ
- GPTQ
- GGUF
Generic deployment commands shown by third-party applications should not be assumed compatible. Use the tested Optimum-Quanto loading method above unless the application specifically documents support for persistent Optimum-Quanto models.
Validation
The finished checkpoint was:
- reloaded with
QuantizedModelForCausalLM.from_pretrained - explicitly transferred to CUDA
- verified to retain Quanto
QLinear/WeightQBytesTensorweights for an internal projection andlm_head - successfully used for CUDA text generation
- checked for tokenizer equivalence with the source Qwen3-4B tokenizer
- checked with the previous tokenizer-regex warning absent
- checked with the previous pad-token fallback warning absent
Observed during a short standalone validation:
- CUDA allocation after transfer: approximately 4.50 GiB
- PyTorch peak allocated: approximately 5.24 GiB
- PyTorch peak reserved: approximately 5.33 GiB
- Windows Task Manager peak: approximately 6.5 GB
Actual memory use varies with prompt length, output length, allocator state, CUDA/PyTorch versions, and the host application.
Download statistics
Hugging Face tracks repository downloads server-side. The repository includes
the standard config.json, so no custom tracking script or telemetry is
required.
License and attribution
Derived from Qwen/Qwen3-4B, released under the Apache License 2.0. This repository changes the weight representation for inference and does not claim authorship of the original Qwen3 model.
Original model: https://huggingface.co/Qwen/Qwen3-4B
- Downloads last month
- 19