Sakura — K2-Horizon-0.9B (GGUF)

Sakura logo

K2-Horizon-0.9B is the smallest member of IFM's open K2-Horizon family: a tiny decoder-only reasoning and agent model with a long context window (Apache-2.0, fully open training data and recipe). This repository holds one GGUF file made from the BF16 weights with our measured mixed-codec method. Independent community quantization, not an official IFM (MBZUAI Institute of Foundation Models) release. Part of the Sakura Mini line.

Runs in mainline llama.cpp. K2 Horizon support was merged on 6 October 2026 (#29535, closes feature request #29424); use a build from b11454 on. Older builds need IFM's branch model/K2Horizon of MBZUAI-IFM/llama.cpp. We loaded this file and generated text with upstream commit 9c2e0e4 (Vulkan).

Quality at a glance

Same method, same texts, same machine for every file; lower KL divergence means closer to the BF16 original:

  • Sakura-K2-Horizon-0.9B-0.43GiB.gguf vs llama.cpp default IQ3_XXS (0.42 GiB): KL divergence to the BF16 original 6 % lower on average over all 3 texts (0.01 GiB larger).

Why: the bit budget per tensor is allocated from measured sensitivity (see How it was made), not from fixed rules. Full table below.

Which file should I take?

File For
Sakura-K2-Horizon-0.9B-0.43GiB.gguf (0.43 GiB, 3.39 bpw) the file we publish for this model: better than the standard reference of the same size

The file

File Size bits/weight KLD en KLD dev KLD wiki same top token (avg) Tensor types by size
Sakura-K2-Horizon-0.9B-0.43GiB.gguf 0.43 GiB 3.39 0.1966 0.1465 0.3266 81.1 % IQ3_XXS 54%, IQ4_XS 11%, Q8_0 10%, IQ2_S 9%, Q4_K 9%, Q2_K 5%

"Tensor types by size" lists the share of the file's bytes per ggml type (all are standard llama.cpp types, so the files run in stock llama.cpp). Quality columns: mean KL divergence of the quantized model's next-token distribution against the BF16 source on held-out text (lower is better) and how often the most likely token is unchanged (higher is better).

Comparison at similar size

Same texts, same context, same chunks, measured by us with llama-perplexity (BF16 source = reference; PPL of the BF16 model: en: 2.029, dev: 1.893, wiki: 15.449). The reference quants are: llama.cpp default recipe, same importance matrix (ours), measured the same way. Sorted by size.

Quant Source Size KLD en KLD dev KLD wiki PPL en same top token (avg)
IQ3_XXS llama.cpp default recipe, same importance matrix (ours) 0.42 GiB 0.2215 0.1670 0.3065 2.519 81.2 %
Sakura-K2-Horizon-0.9B-0.43GiB.gguf this repository 0.43 GiB 0.1966 0.1465 0.3266 2.418 81.1 %
IQ4_XS llama.cpp default recipe, same importance matrix (ours) 0.57 GiB 0.0293 0.0256 0.0483 2.067 92.4 %
IFM official Q4_K_M IFM (official) 0.62 GiB 0.0847 0.0732 0.1077 2.166 88.1 %
IFM official Q5_K_M IFM (official) 0.72 GiB 0.0219 0.0202 0.0288 2.078 93.0 %

Method: 12 chunks of 512 tokens per text; texts: English held-out text, developer-text held-out set, English encyclopedia held-out text (not used in any calibration). Single measurement on one machine; small KLD differences at similar size are not a quality ranking.

How it was made

  1. The BF16 GGUF is IFM's official file from IFM/K2-Horizon-0.9B-GGUF (unchanged); the importance matrix was computed by us (llama-imatrix).
  2. Every large weight matrix was quantized once per candidate type (Q2_K, IQ2_S, IQ3_XXS, IQ3_S, Q3_K, IQ4_XS, Q4_K, Q5_K, Q6_K, Q8_0) with llama-quantize and the importance matrix; the error of each choice was estimated per matrix (importance-weighted) and scaled by a sensitivity factor measured with real KL divergence: every group of tensors (for example the FFN down projections, the attention value projections, the first and last layers) was moved alone to a lower-bit type, and the KL divergence it caused was compared with its quantization error.
  3. An exact budget allocation (multiple-choice knapsack over bytes) picked one type per matrix for each target size; the final model was assembled from the stored tensors without re-quantizing. We do not publish per-tensor choices. Norms, small tensors and the multi-token-prediction layers stay at high precision.

Limits

  • Measured only with KL divergence and perplexity on short held-out texts and not with downstream benchmarks; do not read it as a task-quality claim.
  • Quantization always costs some quality; the smallest file costs the most.
  • llama.cpp version: the K2Horizon architecture is in mainline llama.cpp from build b11454 (pull request #29535, merged 6 October 2026); older builds need IFM's branch model/K2Horizon of MBZUAI-IFM/llama.cpp. The quality numbers on this card were measured with that IFM branch (Vulkan backend; CPU and Vulkan agree to 0.2 % perplexity on our check); we did not re-measure with the mainline build, only checked that the file loads and generates.
  • The model uses the <|ifm|im_start|> chat format with a thinking block; the chat template is embedded in the GGUF.

Credits


中文说明 · 樱花 (Simplified Chinese)

English above. 本节为上文的中文翻译(Sakura = 樱花 yīnghuā);完整的独立中文版见 README_zh.md。

Sakura — K2-Horizon-0.9B(GGUF)

Sakura logo

K2-Horizon-0.9B 是 IFM 开源 K2-Horizon 系列中最小的成员:一个体积很小的、仅解码器的推理与智能体模型,带有长上下文窗口(Apache-2.0,训练数据和配方完全开放)。本仓库包含 一个 GGUF 文件,由 BF16 权重通过我们实测得出的混合编码方法制作。 这是独立的社区量化版本,不是 IFM(MBZUAI Institute of Foundation Models)的官方发布。属于 Sakura Mini 系列。

可在 llama.cpp 主线中运行。 K2 Horizon 支持已于 2026 年 10 月 6 日合并(#29535,关闭了功能请求 #29424);请使用 b11454 及之后的构建。更旧的构建需要 IFM 在 MBZUAI-IFM/llama.cpp 的 model/K2Horizon 分支。我们用上游提交 9c2e0e4(Vulkan)加载了该文件并生成了文本。

质量概览

每个文件使用相同的方法、相同的文本、同一台机器;KL 散度越低,越接近 BF16 原始模型:

  • Sakura-K2-Horizon-0.9B-0.43GiB.gguf 对比 llama.cpp 默认的 IQ3_XXS (0.42 GiB):相对于 BF16 原始模型,在全部 3 个文本上平均 KL 散度 低 6 %(大 0.01 GiB)。

原因:每个张量的比特预算是根据实测的敏感度分配的(见 制作方式),而不是按固定规则。完整表格见下文。

我该选哪个文件?

文件 适用场景
Sakura-K2-Horizon-0.9B-0.43GiB.gguf(0.43 GiB,3.39 bpw) 我们为该模型发布的文件:优于相同大小的标准参照

该文件

文件 大小 bits/weight KLD en KLD dev KLD wiki same top token (avg) 各张量类型(按大小)
Sakura-K2-Horizon-0.9B-0.43GiB.gguf 0.43 GiB 3.39 0.1966 0.1465 0.3266 81.1 % IQ3_XXS 54%, IQ4_XS 11%, Q8_0 10%, IQ2_S 9%, Q4_K 9%, Q2_K 5%

“各张量类型(按大小)”列出每个 ggml 类型占文件字节数的比例(全部是标准的 llama.cpp 类型,因此这些文件可在原版 llama.cpp 中运行)。质量列:量化模型的下一个 token 分布相对于 BF16 源文件在留出文本上的平均 KL 散度(越低越好),以及最可能的 token 保持不变的频率(越高越好)。

相近大小的比较

相同的文本、相同的上下文、相同的分块,由我们使用 llama-perplexity 测量(BF16 源文件 = 参照;BF16 模型的 PPL:en: 2.029,dev: 1.893,wiki: 15.449)。参照量化为:llama.cpp 默认配方、相同的重要性矩阵(我们的)、以相同方式测量。按大小排序。

量化 来源 大小 KLD en KLD dev KLD wiki PPL en same top token (avg)
IQ3_XXS llama.cpp 默认配方,相同的重要性矩阵(我们的) 0.42 GiB 0.2215 0.1670 0.3065 2.519 81.2 %
Sakura-K2-Horizon-0.9B-0.43GiB.gguf 本仓库 0.43 GiB 0.1966 0.1465 0.3266 2.418 81.1 %
IQ4_XS llama.cpp 默认配方,相同的重要性矩阵(我们的) 0.57 GiB 0.0293 0.0256 0.0483 2.067 92.4 %
IFM official Q4_K_M IFM(官方) 0.62 GiB 0.0847 0.0732 0.1077 2.166 88.1 %
IFM official Q5_K_M IFM(官方) 0.72 GiB 0.0219 0.0202 0.0288 2.078 93.0 %

方法:每个文本取 12 个块,每块 512 个 token;文本为:英文留出文本、开发者文本留出集、英文百科留出文本(未用于任何校准)。这是在一台机器上的单次测量;相近大小下 KLD 的细小差异并不构成质量排名。

制作方式

  1. BF16 GGUF 是 IFM 在 IFM/K2-Horizon-0.9B-GGUF 提供的官方文件(未改动);重要性矩阵由我们自己(llama-imatrix)计算。
  2. 每个大型权重矩阵针对每种候选类型(Q2_K、IQ2_S、IQ3_XXS、IQ3_S、Q3_K、IQ4_XS、Q4_K、Q5_K、Q6_K、Q8_0)使用 llama-quantize 和重要性矩阵各量化一次;每种选择的误差按矩阵估计(按重要性加权),并乘以一个 用真实 KL 散度测得的敏感度系数:每组张量(例如 FFN down 投影、注意力 value 投影、首尾几层)被单独移到较低比特的类型,并将它造成的 KL 散度与其量化误差进行比较。
  3. 精确的预算分配(以字节为单位的多选背包问题)为每个目标大小为每个矩阵挑选一种类型;最终模型由已存储的张量组装而成,无需重新量化。 我们不公布逐张量的选择。归一化层、小张量和多 token 预测层保持高精度。

局限

  • 仅用短的留出文本上的 KL 散度和困惑度测量,没有使用下游基准测试;请勿将其理解为任务质量的声明。
  • 量化总会损失一些质量;最小的文件损失最大。
  • llama.cpp 版本: K2Horizon 架构自构建 b11454 起进入 llama.cpp 主线(拉取请求 #29535,于 2026 年 10 月 6 日合并);更旧的构建需要 IFM 在 MBZUAI-IFM/llama.cpp 的 model/K2Horizon 分支。本卡片上的质量数据是用该 IFM 分支测得的(Vulkan 后端;在我们的检查中 CPU 与 Vulkan 的困惑度相差 0.2 %);我们没有用主线构建重新测量,只检查了该文件能够加载并生成文本。
  • 该模型使用带有思考块的 <|ifm|im_start|> 聊天格式;聊天模板已嵌入 GGUF。

致谢

Downloads last month
614
GGUF
Model size
1B params
Architecture
k2-horizon
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for webmp3/Sakura-K2-Horizon-0.9B-GGUF

Quantized
(13)
this model

Collection including webmp3/Sakura-K2-Horizon-0.9B-GGUF