SmoothQuant ibm-granite/granite-3.3-2b-instruct

Quantized derivative of ibm-granite/granite-3.3-2b-instruct. Not Hub fp16. Weights are the packed/runtime artifact from job 20260920T012426Z-d2f5c6.

WikiText-2 (test, 2048-token windows, 65504 tokens)

Measured on NVIDIA A10G (g5.2xlarge) against the original fp16 snapshot. Numbers copied from jobs/20260920T012426Z-d2f5c6/benchmark.json.

Quantized fp16 snapshot
Perplexity 7.762474 7.715887
NLL loss 2.049301 2.043281
Tokens/s (2048 prefill, after warmup) 9501.7 7184.6
Peak VRAM (GB) 3.377 5.641
  • ppl_ratio: 1.006038 (quality_ok=True)
  • improved_throughput: True
  • improved_vram: True

Pull

from huggingface_hub import snapshot_download
path = snapshot_download("chris320211/smoothquant-granite33-2b-g52xlarge")

This is not a drop-in AutoModelForCausalLM.from_pretrained checkpoint. Reload with the included quant_agent_inference_adapter.py after the method repo (and overlay, if any) is on QUANT_AGENT_METHOD_REPO. See quantization_config.json.

License

Base model license (MIT for Phi-3) plus the method repository license. Keep LICENSE / NOTICE.md from the snapshot when present.

Downloads last month
297
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for chris320211/smoothquant-granite33-2b-g52xlarge

Finetuned
(29)
this model

Collection including chris320211/smoothquant-granite33-2b-g52xlarge