This is 2bpw, 8x RTX 6000 (or 6~8x DGX Spark), quant of Kimi K3, with a KLD of only 0.08 and surprisingly strong coding/benchmark performance.

You'll need a vLLM based on: https://github.com/local-inference-lab/vllm/tree/dev/infernal-invocation (docker container soon) b12x for the special kernels required: https://github.com/local-inference-lab/b12x

More details/instructions to follow, but technical information can be found here: https://github.com/local-inference-lab/qsrt/blob/master/docs/qsrt-2bpw-codec.md

Downloads last month
15
Safetensors
Model size
57B params
Tensor type
F32
BF16
F8_E4M3
Inference Providers NEW
This model isn't deployed by any Inference Provider. 馃檵 Ask for provider support

Model tree for lukealonso/Kimi-K3-QSRT-K2

Finetuned
(42)
this model