OSTQuant Qwen3-4B W4A4KV16 (INT4-packed, compact)

Bit-packed to INT4 from the raw OSTQuant W4A4KV16 checkpoint. Cannot be loaded via vanilla AutoModelForCausalLM.from_pretrained(...) — the OSTQuant runtime is required to unpack and evaluate correctly.

  • Compressed from ~8.9 GB (raw .bin) to ~2.9 GB (3× compression).
  • Weights on 4-bit grid. Activations also quantized to 4 bits at inference.
  • Reference PPL when loaded via OSTQuant runtime: 16.49 (WikiText-2 seq 2048, raw .bin).

Files

  • model.safetensors - packed int4 weights + scales + zeros
  • config.json - HF config with ostquant_int4_packed metadata
  • Tokenizer files copied from Qwen/Qwen3-4B

Source

  • Raw: TrojAI/ostquant_qwen3_4b_w4a4kv16_raw
  • Packing script: rebuild_int4_from_raw.py
Downloads last month
-
Safetensors
Model size
1.0B params
Tensor type
I32
·
BF16
·
U8
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for TrojAI/ostquant_qwen3_4b_w4a4kv16_int4_v2

Finetuned
Qwen/Qwen3-4B
Finetuned
(1064)
this model