PrismQuant-Qwen3-Base
Paper · Code · Loading guide
Official checkpoints for PrismQuant: Optimal Null-Space Rotations for Grouped Quantizers. PrismQuant aligns dominant activation directions with the constant directions of asymmetric quantization groups.
Models
| Model key | Base model | Quantization |
|---|---|---|
qwen3_0.6b_base |
Qwen/Qwen3-0.6B-Base | W4A4KV4 |
qwen3_1.7b_base |
Qwen/Qwen3-1.7B-Base | W4A4KV4 |
qwen3_4b_base |
Qwen/Qwen3-4B-Base | W4A4KV4 |
qwen3_8b_base |
Qwen/Qwen3-8B-Base | W4A4KV4 |
Quick start
git clone --depth 1 https://github.com/ForeverBlue816/PrismQuant.git
cd PrismQuant
pip install -e .
from prismquant import load_model
loaded = load_model("qwen3_0.6b_base")
print(loaded.generate("The key idea behind quantization is", max_new_tokens=64))
The loader selects the default PrismQuant checkpoint, downloads its required files at a pinned revision, and obtains the matching base model and tokenizer. These are base models for text completion.
Checkpoint format
The GPTQ INT4 values are stored as dequantized floating-point tensors. The reference runtime applies activation and KV quantization while retaining floating-point storage. Use the PrismQuant loader; this repository is not a standalone Transformers checkpoint or a packed INT4 serving model.
Reference compute: float32. Budget for the base-model state, activations and KV cache as well as the downloaded tensors.
config.json declares the PrismQuant artifact format and default models. checkpoints/ contains decoder weights; rotations/ contains the factors required by the loader.
License
Qwen-derived weights retain Apache-2.0.
- Downloads last month
- 358
Model tree for ForeverBlue/PrismQuant-Qwen3-Base
Base model
Qwen/Qwen3-0.6B-Base