You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Inkling-Small-MLX-4bit

Built with Inkling (Thinking Machines Lab).

MLX (Apple Silicon) conversion of thinkingmachines/Inkling-Small, quantized to 4-bit (affine group quant, group size 64).

Code / loader: github.com/PipeNetwork/inkling-mlx

Inkling Small is a 276B-total / 12B-active sparse-MoE, natively multimodal model (text + image/video + audio → text). This is the full multimodal conversion: all three towers (text backbone, HMLP vision, dMel audio) are ported; the multi-token-prediction head is dropped (inference-irrelevant).

Builds

Variant Size Text ppl Notes
8bit ~280 GB 5.569 near-lossless
6bit ~214 GB 5.569 high quality
4bit ~148 GB 5.452 balanced default
3bit ~115 GB 6.706 ⚠️ experimental — visibly degraded

Perplexity is teacher-forcing over one fixed held-out set (prose / code / reasoning / multilingual) — identical inputs across builds, so the columns compare directly. 4-bit shows no measurable loss vs 8-bit.

There is also a REAP-pruned build: REAP25-4bit keeps 4-bit precision with 192 of 256 routed experts, fitting a 128 GB Mac at ~112 GB for no measurable perplexity cost, with vision and speech intact.

No bf16 build is published. The MLX bf16 conversion is bit-identical to the upstream checkpoint (name-mapping and layout only — the dtype cast is a no-op), so it would carry nothing thinkingmachines/Inkling-Small does not already have, and at ~527 GB it does not fit a 512 GB Mac. If you want it as a requant source, scripts/convert_all.sh regenerates it from the upstream weights in about three minutes.

Quantization scheme: affine int4 (not NVFP4 / MXFP4)

MLX supports FP4 modes and Thinking Machines ships an Inkling-NVFP4 checkpoint — so for the record, we benchmarked round-trip reconstruction error (‖W − Ŵ‖ / ‖W‖ vs bf16) on real Inkling expert weights:

Scheme bits/weight reconstruction error
affine int4 (group 64) 4.50 ~9.1%
nvfp4 (group 16) 4.50 ~10.2%
mxfp4 (group 32) 4.25 ~12.3%

Affine int4 is the most faithful: it is asymmetric (per-group scale and zero-point, 16 uniform levels), which centers on Inkling's near-Gaussian expert weights better than symmetric FP4's fixed non-uniform levels. FP4's real payoff is heavy-tailed activations and native Blackwell FP4 tensor cores — neither helps weight fidelity on Apple Silicon, where MLX would dequantize FP4 anyway. So these builds use affine int4.

⚠️ Loading requires the bundled inkling_mlx loader

The inkling_mm_model architecture is not in stock mlx-lm / mlx-vlm, so this repo bundles a minimal, numerically-validated MLX implementation under inkling_mlx/.

pip install mlx mlx-lm transformers
from inkling_mlx.load import load
from inkling_mlx.generate import greedy_generate
from transformers import AutoTokenizer

model, config = load("/path/to/this/repo")
tok = AutoTokenizer.from_pretrained("/path/to/this/repo", trust_remote_code=True)
ids = tok("The capital of France is")["input_ids"]
print(tok.decode(greedy_generate(model, config, ids, max_new_tokens=64)))

Needs an Apple-Silicon Mac with enough unified memory to hold the weights (≈ the size above).

Status & caveats

  • Text generation works end-to-end via an incremental KV + short-convolution cache.
  • Multimodal is supported end-to-end: the vision/audio towers and their preprocessing (InklingProcessor — image patchify/normalize, audio log-mel→dMel, validated ~1e-7 vs the reference) are included. Pass images/audio via the processor.
  • Quantized: attention / MLP / expert projections, token embed+unembed, and the vision/audio matmuls. Kept in higher precision: the MoE router, RMSNorms, the four short-convolutions per layer, and the relative-position bias.

Conversion is streaming (tensor-by-tensor; the ~527 GB bf16 model never fully loads into RAM) and was validated with fp32 numerical parity against transformers PR #47347. License: Apache-2.0 (inherits the base model).

Downloads last month
4
Safetensors
Model size
41B params
Tensor type
BF16
·
U32
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bielquant/Inkling-Small-MLX-4bit

Quantized
(37)
this model