Inkling-Small EXL3 quantization suite

This is a lightweight index card for the public EXL3 quantizations of thinkingmachines/Inkling-Small. It intentionally contains no model weights.

Hugging Face Collection

Browse the complete suite in the public Inkling-Small EXL3 Quantization Suite collection.

Variants

The target bpw applies to routed-expert trellis weights. Attention, shared experts, multimodal components, the LM head, and MTP remain in source precision, so complete repository sizes are larger than a whole-model quantization at the nominal rate. Each variant card documents its layout, fractional allocation, validation status, and compatibility limits.

Runtime status: All seven archives are structurally verified and downloadable. Full-model text generation, multimodal generation, and MTP validation remain pending; these repositories are not currently drop-in Transformers checkpoints.

Credits

Thanks to the Thinking Machines Lab Inkling team for Inkling-Small, TurboDerp for ExLlamaV3/EXL3, and JarvisLabs for the 8× NVIDIA H200 compute used for the sweep.

This is an independent community quantization and is not an official release from Thinking Machines Lab, TurboDerp, or JarvisLabs.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support