OSTQuant Qwen3-4B W4A4KV16 (INT4-packed, compact)
Bit-packed to INT4 from the raw OSTQuant W4A4KV16 checkpoint. Cannot be loaded
via vanilla AutoModelForCausalLM.from_pretrained(...) — the OSTQuant runtime
is required to unpack and evaluate correctly.
- Compressed from ~8.9 GB (raw .bin) to ~2.9 GB (3× compression).
- Weights on 4-bit grid. Activations also quantized to 4 bits at inference.
- Reference PPL when loaded via OSTQuant runtime: 16.49 (WikiText-2 seq 2048, raw .bin).
Files
model.safetensors- packed int4 weights + scales + zerosconfig.json- HF config withostquant_int4_packedmetadata- Tokenizer files copied from Qwen/Qwen3-4B
Source
- Raw: TrojAI/ostquant_qwen3_4b_w4a4kv16_raw
- Packing script: rebuild_int4_from_raw.py
- Downloads last month
- -
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support