Instructions to use webmp3/Sakura-K2-Horizon-0.9B-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use webmp3/Sakura-K2-Horizon-0.9B-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf webmp3/Sakura-K2-Horizon-0.9B-GGUF # Run inference directly in the terminal: llama cli -hf webmp3/Sakura-K2-Horizon-0.9B-GGUF
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf webmp3/Sakura-K2-Horizon-0.9B-GGUF # Run inference directly in the terminal: llama cli -hf webmp3/Sakura-K2-Horizon-0.9B-GGUF
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf webmp3/Sakura-K2-Horizon-0.9B-GGUF # Run inference directly in the terminal: ./llama-cli -hf webmp3/Sakura-K2-Horizon-0.9B-GGUF
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf webmp3/Sakura-K2-Horizon-0.9B-GGUF # Run inference directly in the terminal: ./build/bin/llama-cli -hf webmp3/Sakura-K2-Horizon-0.9B-GGUF
Use Docker
docker model run hf.co/webmp3/Sakura-K2-Horizon-0.9B-GGUF
- LM Studio
- Jan
- vLLM
How to use webmp3/Sakura-K2-Horizon-0.9B-GGUF with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "webmp3/Sakura-K2-Horizon-0.9B-GGUF" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "webmp3/Sakura-K2-Horizon-0.9B-GGUF", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/webmp3/Sakura-K2-Horizon-0.9B-GGUF
- Ollama
How to use webmp3/Sakura-K2-Horizon-0.9B-GGUF with Ollama:
ollama run hf.co/webmp3/Sakura-K2-Horizon-0.9B-GGUF
- Unsloth Desktop
- Pi
How to use webmp3/Sakura-K2-Horizon-0.9B-GGUF with Pi:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webmp3/Sakura-K2-Horizon-0.9B-GGUF
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "llama-cpp": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "webmp3/Sakura-K2-Horizon-0.9B-GGUF" } ] } } }Run Pi
# Start Pi in your project directory: pi
- Docker Model Runner
How to use webmp3/Sakura-K2-Horizon-0.9B-GGUF with Docker Model Runner:
docker model run hf.co/webmp3/Sakura-K2-Horizon-0.9B-GGUF
- Lemonade
How to use webmp3/Sakura-K2-Horizon-0.9B-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull webmp3/Sakura-K2-Horizon-0.9B-GGUF
Run and chat with the model
lemonade run user.Sakura-K2-Horizon-0.9B-GGUF-{{QUANT_TAG}}List all available models
lemonade list
- Hermes Agent
How to use webmp3/Sakura-K2-Horizon-0.9B-GGUF with Hermes Agent:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webmp3/Sakura-K2-Horizon-0.9B-GGUF
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default webmp3/Sakura-K2-Horizon-0.9B-GGUF
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use webmp3/Sakura-K2-Horizon-0.9B-GGUF with OpenClaw:
Start the llama.cpp server
# Install llama.cpp: brew install llama.cpp # Start a local OpenAI-compatible server: llama serve -hf webmp3/Sakura-K2-Horizon-0.9B-GGUF
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "webmp3/Sakura-K2-Horizon-0.9B-GGUF" \ --custom-provider-id llama-cpp \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
Sakura — K2-Horizon-0.9B (GGUF)
K2-Horizon-0.9B is the smallest member of IFM's open K2-Horizon family: a tiny decoder-only reasoning and agent model with a long context window (Apache-2.0, fully open training data and recipe). This repository holds one GGUF file made from the BF16 weights with our measured mixed-codec method. Independent community quantization, not an official IFM (MBZUAI Institute of Foundation Models) release. Part of the Sakura Mini line.
Runs in mainline llama.cpp. K2 Horizon support was merged on 6 October 2026 (#29535, closes feature request #29424); use a build from
b11454on. Older builds need IFM's branchmodel/K2Horizonof MBZUAI-IFM/llama.cpp. We loaded this file and generated text with upstream commit9c2e0e4(Vulkan).
Quality at a glance
Same method, same texts, same machine for every file; lower KL divergence means closer to the BF16 original:
Sakura-K2-Horizon-0.9B-0.43GiB.ggufvs llama.cpp default IQ3_XXS (0.42 GiB): KL divergence to the BF16 original 6 % lower on average over all 3 texts (0.01 GiB larger).
Why: the bit budget per tensor is allocated from measured sensitivity (see How it was made), not from fixed rules. Full table below.
Which file should I take?
| File | For |
|---|---|
Sakura-K2-Horizon-0.9B-0.43GiB.gguf (0.43 GiB, 3.39 bpw) |
the file we publish for this model: better than the standard reference of the same size |
The file
| File | Size | bits/weight | KLD en | KLD dev | KLD wiki | same top token (avg) | Tensor types by size |
|---|---|---|---|---|---|---|---|
Sakura-K2-Horizon-0.9B-0.43GiB.gguf |
0.43 GiB | 3.39 | 0.1966 | 0.1465 | 0.3266 | 81.1 % | IQ3_XXS 54%, IQ4_XS 11%, Q8_0 10%, IQ2_S 9%, Q4_K 9%, Q2_K 5% |
"Tensor types by size" lists the share of the file's bytes per ggml type (all are standard llama.cpp types, so the files run in stock llama.cpp). Quality columns: mean KL divergence of the quantized model's next-token distribution against the BF16 source on held-out text (lower is better) and how often the most likely token is unchanged (higher is better).
Comparison at similar size
Same texts, same context, same chunks, measured by us with llama-perplexity (BF16 source = reference; PPL of the BF16 model: en: 2.029, dev: 1.893, wiki: 15.449). The reference quants are: llama.cpp default recipe, same importance matrix (ours), measured the same way. Sorted by size.
| Quant | Source | Size | KLD en | KLD dev | KLD wiki | PPL en | same top token (avg) |
|---|---|---|---|---|---|---|---|
| IQ3_XXS | llama.cpp default recipe, same importance matrix (ours) | 0.42 GiB | 0.2215 | 0.1670 | 0.3065 | 2.519 | 81.2 % |
Sakura-K2-Horizon-0.9B-0.43GiB.gguf |
this repository | 0.43 GiB | 0.1966 | 0.1465 | 0.3266 | 2.418 | 81.1 % |
| IQ4_XS | llama.cpp default recipe, same importance matrix (ours) | 0.57 GiB | 0.0293 | 0.0256 | 0.0483 | 2.067 | 92.4 % |
| IFM official Q4_K_M | IFM (official) | 0.62 GiB | 0.0847 | 0.0732 | 0.1077 | 2.166 | 88.1 % |
| IFM official Q5_K_M | IFM (official) | 0.72 GiB | 0.0219 | 0.0202 | 0.0288 | 2.078 | 93.0 % |
Method: 12 chunks of 512 tokens per text; texts: English held-out text, developer-text held-out set, English encyclopedia held-out text (not used in any calibration). Single measurement on one machine; small KLD differences at similar size are not a quality ranking.
How it was made
- The BF16 GGUF is IFM's official file from IFM/K2-Horizon-0.9B-GGUF (unchanged); the importance matrix was computed by us (
llama-imatrix). - Every large weight matrix was quantized once per candidate type (Q2_K, IQ2_S, IQ3_XXS, IQ3_S, Q3_K, IQ4_XS, Q4_K, Q5_K, Q6_K, Q8_0) with
llama-quantizeand the importance matrix; the error of each choice was estimated per matrix (importance-weighted) and scaled by a sensitivity factor measured with real KL divergence: every group of tensors (for example the FFN down projections, the attention value projections, the first and last layers) was moved alone to a lower-bit type, and the KL divergence it caused was compared with its quantization error. - An exact budget allocation (multiple-choice knapsack over bytes) picked one type per matrix for each target size; the final model was assembled from the stored tensors without re-quantizing. We do not publish per-tensor choices. Norms, small tensors and the multi-token-prediction layers stay at high precision.
Limits
- Measured only with KL divergence and perplexity on short held-out texts and not with downstream benchmarks; do not read it as a task-quality claim.
- Quantization always costs some quality; the smallest file costs the most.
- llama.cpp version: the
K2Horizonarchitecture is in mainline llama.cpp from buildb11454(pull request #29535, merged 6 October 2026); older builds need IFM's branchmodel/K2Horizonof MBZUAI-IFM/llama.cpp. The quality numbers on this card were measured with that IFM branch (Vulkan backend; CPU and Vulkan agree to 0.2 % perplexity on our check); we did not re-measure with the mainline build, only checked that the file loads and generates. - The model uses the
<|ifm|im_start|>chat format with a thinking block; the chat template is embedded in the GGUF.
Credits
- Original model: K2-Horizon-0.9B (IFM / MBZUAI, Apache-2.0), Apache-2.0 license (included as
LICENSE). - BF16 GGUF and the official Q4_K_M / Q5_K_M reference quants: IFM/K2-Horizon-0.9B-GGUF. The other reference quants are
llama-quantizedefault recipes with our importance matrix. - GGUF format and tools: ggml-org/llama.cpp (MIT).
中文说明 · 樱花 (Simplified Chinese)
English above. 本节为上文的中文翻译(Sakura = 樱花 yīnghuā);完整的独立中文版见 README_zh.md。
Sakura — K2-Horizon-0.9B(GGUF)
K2-Horizon-0.9B 是 IFM 开源 K2-Horizon 系列中最小的成员:一个体积很小的、仅解码器的推理与智能体模型,带有长上下文窗口(Apache-2.0,训练数据和配方完全开放)。本仓库包含 一个 GGUF 文件,由 BF16 权重通过我们实测得出的混合编码方法制作。 这是独立的社区量化版本,不是 IFM(MBZUAI Institute of Foundation Models)的官方发布。属于 Sakura Mini 系列。
可在 llama.cpp 主线中运行。 K2 Horizon 支持已于 2026 年 10 月 6 日合并(#29535,关闭了功能请求 #29424);请使用
b11454及之后的构建。更旧的构建需要 IFM 在 MBZUAI-IFM/llama.cpp 的model/K2Horizon分支。我们用上游提交9c2e0e4(Vulkan)加载了该文件并生成了文本。
质量概览
每个文件使用相同的方法、相同的文本、同一台机器;KL 散度越低,越接近 BF16 原始模型:
Sakura-K2-Horizon-0.9B-0.43GiB.gguf对比 llama.cpp 默认的 IQ3_XXS (0.42 GiB):相对于 BF16 原始模型,在全部 3 个文本上平均 KL 散度 低 6 %(大 0.01 GiB)。
原因:每个张量的比特预算是根据实测的敏感度分配的(见 制作方式),而不是按固定规则。完整表格见下文。
我该选哪个文件?
| 文件 | 适用场景 |
|---|---|
Sakura-K2-Horizon-0.9B-0.43GiB.gguf(0.43 GiB,3.39 bpw) |
我们为该模型发布的文件:优于相同大小的标准参照 |
该文件
| 文件 | 大小 | bits/weight | KLD en | KLD dev | KLD wiki | same top token (avg) | 各张量类型(按大小) |
|---|---|---|---|---|---|---|---|
Sakura-K2-Horizon-0.9B-0.43GiB.gguf |
0.43 GiB | 3.39 | 0.1966 | 0.1465 | 0.3266 | 81.1 % | IQ3_XXS 54%, IQ4_XS 11%, Q8_0 10%, IQ2_S 9%, Q4_K 9%, Q2_K 5% |
“各张量类型(按大小)”列出每个 ggml 类型占文件字节数的比例(全部是标准的 llama.cpp 类型,因此这些文件可在原版 llama.cpp 中运行)。质量列:量化模型的下一个 token 分布相对于 BF16 源文件在留出文本上的平均 KL 散度(越低越好),以及最可能的 token 保持不变的频率(越高越好)。
相近大小的比较
相同的文本、相同的上下文、相同的分块,由我们使用 llama-perplexity 测量(BF16 源文件 = 参照;BF16 模型的 PPL:en: 2.029,dev: 1.893,wiki: 15.449)。参照量化为:llama.cpp 默认配方、相同的重要性矩阵(我们的)、以相同方式测量。按大小排序。
| 量化 | 来源 | 大小 | KLD en | KLD dev | KLD wiki | PPL en | same top token (avg) |
|---|---|---|---|---|---|---|---|
| IQ3_XXS | llama.cpp 默认配方,相同的重要性矩阵(我们的) | 0.42 GiB | 0.2215 | 0.1670 | 0.3065 | 2.519 | 81.2 % |
Sakura-K2-Horizon-0.9B-0.43GiB.gguf |
本仓库 | 0.43 GiB | 0.1966 | 0.1465 | 0.3266 | 2.418 | 81.1 % |
| IQ4_XS | llama.cpp 默认配方,相同的重要性矩阵(我们的) | 0.57 GiB | 0.0293 | 0.0256 | 0.0483 | 2.067 | 92.4 % |
| IFM official Q4_K_M | IFM(官方) | 0.62 GiB | 0.0847 | 0.0732 | 0.1077 | 2.166 | 88.1 % |
| IFM official Q5_K_M | IFM(官方) | 0.72 GiB | 0.0219 | 0.0202 | 0.0288 | 2.078 | 93.0 % |
方法:每个文本取 12 个块,每块 512 个 token;文本为:英文留出文本、开发者文本留出集、英文百科留出文本(未用于任何校准)。这是在一台机器上的单次测量;相近大小下 KLD 的细小差异并不构成质量排名。
制作方式
- BF16 GGUF 是 IFM 在 IFM/K2-Horizon-0.9B-GGUF 提供的官方文件(未改动);重要性矩阵由我们自己(
llama-imatrix)计算。 - 每个大型权重矩阵针对每种候选类型(Q2_K、IQ2_S、IQ3_XXS、IQ3_S、Q3_K、IQ4_XS、Q4_K、Q5_K、Q6_K、Q8_0)使用
llama-quantize和重要性矩阵各量化一次;每种选择的误差按矩阵估计(按重要性加权),并乘以一个 用真实 KL 散度测得的敏感度系数:每组张量(例如 FFN down 投影、注意力 value 投影、首尾几层)被单独移到较低比特的类型,并将它造成的 KL 散度与其量化误差进行比较。 - 精确的预算分配(以字节为单位的多选背包问题)为每个目标大小为每个矩阵挑选一种类型;最终模型由已存储的张量组装而成,无需重新量化。 我们不公布逐张量的选择。归一化层、小张量和多 token 预测层保持高精度。
局限
- 仅用短的留出文本上的 KL 散度和困惑度测量,没有使用下游基准测试;请勿将其理解为任务质量的声明。
- 量化总会损失一些质量;最小的文件损失最大。
- llama.cpp 版本:
K2Horizon架构自构建b11454起进入 llama.cpp 主线(拉取请求 #29535,于 2026 年 10 月 6 日合并);更旧的构建需要 IFM 在 MBZUAI-IFM/llama.cpp 的model/K2Horizon分支。本卡片上的质量数据是用该 IFM 分支测得的(Vulkan 后端;在我们的检查中 CPU 与 Vulkan 的困惑度相差 0.2 %);我们没有用主线构建重新测量,只检查了该文件能够加载并生成文本。 - 该模型使用带有思考块的
<|ifm|im_start|>聊天格式;聊天模板已嵌入 GGUF。
致谢
- 原始模型:K2-Horizon-0.9B (IFM / MBZUAI, Apache-2.0),Apache-2.0 许可证(以
LICENSE随附)。 - BF16 GGUF 以及官方 Q4_K_M / Q5_K_M 参照量化:IFM/K2-Horizon-0.9B-GGUF。其他参照量化是
llama-quantize的默认配方,配以我们的重要性矩阵。 - GGUF 格式和工具:ggml-org/llama.cpp (MIT)。
- Downloads last month
- 614
We're not able to determine the quantization variants.
Model tree for webmp3/Sakura-K2-Horizon-0.9B-GGUF
Base model
IFM/K2-Horizon-0.9B