Kimi-Linear-48B-A3B-Instruct β€” APEX GGUF

MoE-aware, mixed-precision APEX quantizations of moonshotai/Kimi-Linear-48B-A3B-Instruct β€” 48B total / ~3B active, a hybrid linear-attention MoE: most layers use KDA (Kimi Delta Attention, gated-delta linear attention), a few use full MLA attention, over a 256-routed + 1-shared expert FFN.

To my knowledge this is the first APEX quant of a linear-attention hybrid MoE. APEX assigns precision per tensor role and per layer instead of uniformly; here that meant teaching the recipe about tensor families the stock generator doesn't know (see Method).

Results

Perplexity on wikitext-2-raw (test, 200Γ—512-token windows), llama-perplexity.

File Size BPW PPL Ξ” vs bf16
bf16 (reference) 92 GB 16.0 7.374 β€”
APEX-balanced (no-imatrix) 33 GB 5.72 7.377 +0.04%
APEX-handroll (ssm@Q8_0, no-imatrix) 33 GB 5.72 7.382 +0.11%
APEX-i-quality (imatrix, IQ4_XS mid experts) 29 GB 4.94 pending pending

Both no-imatrix tiers land essentially on the bf16 reference (within ~0.1%). An imatrix-guided i-quality tier (IQ4_XS mid experts, 29 GB) is now available as Kimi-Linear-48B-A3B-Instruct-APEX-i-quality.gguf; its perplexity benchmark is still being run and will be filled in here.

From a 92 GB bf16 baseline β†’ 33 GB (~2.8Γ— smaller), and it runs on a 128 GB unified-memory box (fits with full GPU offload). Coherent on general and factual prompts.

The balanced and handroll tiers were built without an imatrix (Q6_K/Q5_K experts, Q8_0 shared, Q6_K attention β€” none of which require importance data). The newer i-quality tier is imatrix-guided (IQ4_XS mid experts). Deeper lower-bit "I-tier" variants (IQ3/IQ2) would also need an imatrix and are not included here.

Note on the two tiers (a null result)

The hand-roll tier pins the KDA recurrence tensors (ssm_conv1d_*, ssm_f/g_*, ssm_beta) to Q8_0 instead of Q6_K, testing whether protecting the linear-attention state preserves quality. It doesn't β€” PPL is identical within noise (7.382 vs 7.377), at the same size (the ssm tensors are tiny next to the experts). Use balanced. The hand-roll is kept only to document the experiment.

Which file

  • APEX-balanced β€” recommended. Q6_K/Q5_K experts on a layer-depth gradient, Q8_0 shared experts, Q6_K attention + KDA tensors.
  • APEX-i-quality β€” imatrix-guided IQ4_XS mid experts (29 GB); the smallest tier here. Try it when you want to save a few GB over balanced; PPL comparison pending.
  • APEX-handroll β€” experimental (see null-result note below); not recommended.

Usage (llama.cpp)

llama-cli   -m Kimi-Linear-48B-A3B-Instruct-APEX-balanced.gguf -ngl 999 -p "Hello"
llama-server -m Kimi-Linear-48B-A3B-Instruct-APEX-balanced.gguf -ngl 999 --host 0.0.0.0 --port 8080

Requires a llama.cpp build supporting the kimi_linear architecture and the kimi-k2 pre-tokenizer.

Method

APEX is a bit-allocation recipe over stock llama-quantize --tensor-type-file. Kimi-Linear needed two tensor families the stock APEX generator doesn't emit:

  • MLA (full-attention layers): attn_kv_a_mqa, attn_k_b, attn_v_b
  • KDA (linear-attention layers): ssm_conv1d_{k,q,v}, ssm_f_a/f_b, ssm_g_a/g_b, ssm_beta (norms/1-D state kept F32)

plus a dense layer 0 (--dense-layers 1). The expert intermediate dim is 2048 (256-divisible), so no IQ4_NL workaround was needed. Config generation + patching: see REPRODUCE.md, patch_kimi_config.py, and configs/.

Baseline: quantized from bartowski's bf16 GGUF.

Attribution & licenses

All MIT; see LICENSE and NOTICE.

Unofficial community quantization; not affiliated with or endorsed by Moonshot AI.

Downloads last month
690
GGUF
Model size
49B params
Architecture
kimi-linear
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Myric/Kimi-Linear-48B-A3B-Instruct-APEX-GGUF

Quantized
(24)
this model