Qwen3-4B Quanto INT8

INT8 weight-quantized version of Qwen/Qwen3-4B using persistent Optimum-Quanto qint8 weights with BF16 activations.

Important: This is a persistent Optimum-Quanto checkpoint. Load it with optimum.quanto.QuantizedModelForCausalLM as shown below. Do not assume that generic AutoModelForCausalLM, vLLM, SGLang, AWQ, GPTQ, GGUF, or TorchAO loading paths are compatible unless that application explicitly supports persistent Optimum-Quanto checkpoints.

Quantization

  • Base model: Qwen/Qwen3-4B
  • Persistent Optimum-Quanto qint8
  • BF16 activations
  • All Linear layers quantized, including lm_head
  • Repository size: approximately 4.484 GiB
  • No custom trust_remote_code=True loader required
  • PAD: 151643
  • EOS: [151645, 151643]
  • BOS: 151643

Requirements

Validated with:

  • Python 3.12.10
  • PyTorch 2.12.0+cu130
  • Transformers 5.17.0
  • Optimum-Quanto
  • NVIDIA RTX 4060 Ti 16 GB

Install the model-side dependencies:

pip install "transformers>=5.17.0" optimum-quanto safetensors huggingface_hub

Install PyTorch separately using the build appropriate for your CUDA/CPU environment.

TorchAO warning

TorchAO is not required for this model.

If torchao is installed in the same environment, newer PyTorch versions may print deprecation messages mentioning KernelPreference, ScaleCalculationMode, or register_constant(). These originate from the TorchAO/PyTorch dependency combination, not this Qwen3-4B INT8 checkpoint, and do not indicate failed model loading or broken quantization.

If another application requires TorchAO, the messages can be ignored. If TorchAO is not needed, removing it eliminated the messages in the tested environment.

Correct loading method

import torch
from transformers import AutoTokenizer
from optimum.quanto import QuantizedModelForCausalLM

model_id = "Redtash1/Qwen3-4B-INT8"

tokenizer = AutoTokenizer.from_pretrained(model_id)

model = QuantizedModelForCausalLM.from_pretrained(
    model_id,
    dtype=torch.bfloat16,
)

# Explicit transfer is recommended for this persistent Quanto checkpoint.
model = model.to("cuda")
model.eval()

During validation, relying on device_map="cuda" alone did not reliably move the persistent Quanto wrapper to the GPU, so explicit .to("cuda") is recommended.

Example generation

messages = [
    {
        "role": "user",
        "content": "Write a short description of a quiet mountain lake."
    }
]

text = tokenizer.apply_chat_template(
    messages,
    tokenize=False,
    add_generation_prompt=True,
    enable_thinking=False,
)

inputs = tokenizer(text, return_tensors="pt").to("cuda")

with torch.inference_mode():
    output = model.generate(
        **inputs,
        max_new_tokens=256,
        do_sample=True,
        temperature=0.7,
        top_p=0.8,
        top_k=20,
    )

new_tokens = output[0, inputs["input_ids"].shape[1]:]
print(tokenizer.decode(new_tokens, skip_special_tokens=True))

Compatibility

This checkpoint is Optimum-Quanto, not:

  • TorchAO
  • AWQ
  • GPTQ
  • GGUF

Generic deployment commands shown by third-party applications should not be assumed compatible. Use the tested Optimum-Quanto loading method above unless the application specifically documents support for persistent Optimum-Quanto models.

Validation

The finished checkpoint was:

  • reloaded with QuantizedModelForCausalLM.from_pretrained
  • explicitly transferred to CUDA
  • verified to retain Quanto QLinear / WeightQBytesTensor weights for an internal projection and lm_head
  • successfully used for CUDA text generation
  • checked for tokenizer equivalence with the source Qwen3-4B tokenizer
  • checked with the previous tokenizer-regex warning absent
  • checked with the previous pad-token fallback warning absent

Observed during a short standalone validation:

  • CUDA allocation after transfer: approximately 4.50 GiB
  • PyTorch peak allocated: approximately 5.24 GiB
  • PyTorch peak reserved: approximately 5.33 GiB
  • Windows Task Manager peak: approximately 6.5 GB

Actual memory use varies with prompt length, output length, allocator state, CUDA/PyTorch versions, and the host application.

Download statistics

Hugging Face tracks repository downloads server-side. The repository includes the standard config.json, so no custom tracking script or telemetry is required.

License and attribution

Derived from Qwen/Qwen3-4B, released under the Apache License 2.0. This repository changes the weight representation for inference and does not claim authorship of the original Qwen3 model.

Original model: https://huggingface.co/Qwen/Qwen3-4B

Downloads last month
19
Safetensors
Model size
4B params
Tensor type
BF16
·
I8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Redtash1/Qwen3-4B-INT8

Finetuned
Qwen/Qwen3-4B
Finetuned
(1128)
this model