FlatQuant microsoft/Phi-3-mini-4k-instruct

Quantized derivative of microsoft/Phi-3-mini-4k-instruct. Not Hub fp16. Weights are the packed/runtime artifact from job 20260919T202736Z-5a1860.

WikiText-2 (test, 2048-token windows, 65504 tokens)

Measured on NVIDIA A10G (g5.2xlarge) against the original fp16 snapshot. Numbers copied from jobs/20260919T202736Z-5a1860/benchmark.json.

Quantized fp16 snapshot
Perplexity 7.393646 6.332276
NLL loss 2.000621 1.845660
Tokens/s (2048 prefill, after warmup) 5105.9 3387.6
Peak VRAM (GB) 3.355 9.681
  • ppl_ratio: 1.167613 (quality_ok=True)
  • improved_throughput: True
  • improved_vram: True

Pull

from huggingface_hub import snapshot_download
path = snapshot_download("chris320211/flatquant-phi3-mini-4k-g52xlarge")

This is not a drop-in AutoModelForCausalLM.from_pretrained checkpoint. Reload with the included quant_agent_inference_adapter.py after the method repo (and overlay, if any) is on QUANT_AGENT_METHOD_REPO. See quantization_config.json.

License

Base model license (MIT for Phi-3) plus the method repository license. Keep LICENSE / NOTICE.md from the snapshot when present.

Downloads last month
396
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for chris320211/flatquant-phi3-mini-4k-g52xlarge

Finetuned
(1118)
this model

Collection including chris320211/flatquant-phi3-mini-4k-g52xlarge