Sakura — Qwen3.8-Flash-Next builds: overview

Sakura logo

All of our Qwen3.8-Flash-Next GGUF builds in one place: what each one is, how it differs, what we measured, and how to run it. No weights are in this repository. Qwen3.8-Flash-Next (qwen4exp, 512 experts, 10 active, a 51B-parameter per-layer token table that stays on disk) is large; all builds below keep 352 of the 512 experts per layer (chosen on German + English routing) so that they run on a single 64 GB AMD Strix Halo. They are independent community builds, not official Qwen releases. Part of the Sakura Micro collection.

Which one should I take?

Your situation Take
32 GB GPU Swift 2.5-bit (about 26 GiB of GPU weights)
64 GB unified memory (Strix Halo class), best coding result and shortest thinking Swift 3-bit
You want the unmodified base-model weights (no fine-tune) ISTA 3-bit
You want to try Darwin-R3 weights on the same layout ISTA + Darwin-R3 (experimental: in our tests no difference to the ISTA build)

What differs between the builds

Four independent axes: where the weights come from (Swift fine-tune, ISTA base-model quantization, ISTA with Darwin-R3 tensors), bits (3-bit IQ3_XXS or 2.5-bit IQ2_XS), how many experts (all 352 here) and which experts (the same German + English selection, lists below). Naming scheme: Sakura-Qwen3.8-Flash-Next-<source>-<bits>-<experts>E[-<size class>]-GGUF.

Measured results (same runtime, tasks and seeds for every build)

Runtime: Unsloth fork qwen4exp/mtp (commit 6fcaa16, Vulkan, built by us), shared MTP head mtp-Qwen3.8-Flash-Next-shared-Q4_K_M.gguf (n-max 2), -b 2048 -ub 2048, 32k context, KV q8_0, thinking budget 10000, max 12288 tokens, temperature 0.6, top-p 0.95, top-k 20, seeds 1000 + task number, one run per task, own frozen task set of 32 LeetCode/AtCoder-style tasks (grader: unit tests). AMD Strix Halo, 64 GB, Radeon 8060S, Windows, Vulkan.

Build Weights from File Pool at load tasks 1-8 tasks 9-20 tasks 21-32 tokens 9-20 avg decode in tasks 9-20 Note
Swift 3-bit Swift 1.5 fine-tune of Flash-Next 58.2 GiB ~41 GiB (30 dedicated + 10-12 shared) 8/8 11/12 9/12 98,937 28 tok/s recommended when it fits
Swift 2.5-bit (32 GB) Swift 1.5 fine-tune of Flash-Next 53 GiB fits a 32 GB GPU (about 26 GiB of GPU weights) 7/8 10/12 10/12 85,714 33 tok/s for 32 GB cards
ISTA 3-bit ISTA-DASLab GSQ-RCO of the base model (no fine-tune) 58.2 GiB ~41 GiB (30 dedicated + 11 shared) 8/8 11/12 9/12 100,341 28 tok/s base-model weights
ISTA + Darwin-R3 (experimental) ISTA build with 496 Darwin-180B-RSI-R3 tensors grafted in 58.7 GiB ~41 GiB (30 dedicated + 11 shared) 8/8 11/12 10/12 89,952 30 tok/s experimental, no demonstrated advantage

Single runs on one machine: a difference of one task is inside the noise; token counts and decode speed are measured. The Swift 3-bit build also passes a 131072-token context test at -ub 512 (retrieval of five planted facts at 18k to 118k tokens) and a knockout-style coding test (about 99/100). Deep-audit style tests, German probes and perplexity are on the individual model cards.

Runtime notes

  • The shared MTP head only loads in the Unsloth fork (danielhanchen/llama.cpp, branch qwen4exp/mtp, commit 6fcaa16; we built it for Vulkan with MSVC and Vulkan SDK 1.4.357). Official llama.cpp b11259/b11294 run the files without speculation (about 24 tok/s) because the MTP-only head has no loader there.
  • Another community build (agentionai/llama.cpp) uses its own MTP draft file; the two heads are not interchangeable between the two builds.
  • Use -lm dio --lazy-mode on so that the per-layer token table is read lazily. For 131072 context use -ub 512 (-ub 2048 runs out of Vulkan memory there).
  • ROCm on Windows collapses to about 6 tok/s as soon as the resident weights exceed the dedicated GPU memory; the Vulkan numbers above include the spill into shared memory.

Tools and data (this repository)

  • tools/prune_raw_gguf.py: cuts the experts of any qwen4exp GGUF at the byte level (works for every tensor type, including ones unknown to gguf-py).
  • tools/graft_tensors.py: replaces named tensors of a GGUF by tensors fetched with HTTP range requests from another GGUF (used for the Darwin-R3 build); tools/hf_gguf_header.py reads a remote GGUF's tensor table.
  • keep_lists/: kept experts per layer (original numbering) for K320, K336, K352, K368, K384, K416, K448; nested (each smaller set is contained in the next). Ranked by routing mass (hot in German or English) on a German + English calibration mix.

Honest limits

Expert pruning without retraining costs quality; we checked it on German/English text, perplexity and our own coding tasks only. Everything above is single runs. Nothing here is a benchmark suite.


中文说明 · 樱花 (Simplified Chinese)

English above. 本节为上文的中文翻译(Sakura = 樱花 yīnghuā);完整的独立中文版见 README_zh.md。

Sakura — Qwen3.8-Flash-Next 构建版本:概览

Sakura logo

我们所有 Qwen3.8-Flash-Next GGUF 构建版本集中在一处:每个版本是什么、有何不同、我们测量了什么,以及如何运行。本仓库中没有权重。 Qwen3.8-Flash-Next(qwen4exp,512 个专家,激活 10 个,一个 51B 参数的逐层 token 表留在磁盘上)很大;下面所有构建都在每层 512 个专家中保留 352 个(依据德语 + 英语的路由情况选出),以便在单台 64 GB 的 AMD Strix Halo 上运行。 它们是独立的社区构建版本,不是 Qwen 的官方发布。属于 Sakura Micro 合集。

我该选哪一个?

你的情况 选择
32 GB GPU Swift 2.5 比特(GPU 权重约 26 GiB)
64 GB 统一内存(Strix Halo 级别),追求最佳编程结果和最短的思考 Swift 3 比特
你想要未经修改的基础模型权重(未微调) ISTA 3 比特
你想在相同布局上试用 Darwin-R3 权重 ISTA + Darwin-R3(实验性:在我们的测试中与 ISTA 构建没有差别)

各构建之间有何不同

四个相互独立的维度:权重来自哪里(Swift 微调、ISTA 基础模型量化、带 Darwin-R3 张量的 ISTA)、比特数(3 比特 IQ3_XXS 或 2.5 比特 IQ2_XS)、多少个专家(这里都是 352 个)以及 哪些专家(相同的德语 + 英语选择,列表见下文)。 命名方案:Sakura-Qwen3.8-Flash-Next-<source>-<bits>-<experts>E[-<size class>]-GGUF。

测量结果(每个构建使用相同的运行时、任务和随机种子)

运行时:Unsloth 分支 qwen4exp/mtp(提交 6fcaa16,Vulkan,由我们自己构建),共享 MTP 头 mtp-Qwen3.8-Flash-Next-shared-Q4_K_M.gguf(n-max 2),-b 2048 -ub 2048,32k 上下文,KV q8_0,思考预算 10000,最多 12288 个 token,temperature 0.6,top-p 0.95,top-k 20,随机种子 1000 + 任务编号,每个任务运行一次,使用自己的冻结任务集,共 32 个 LeetCode/AtCoder 风格的任务(评分器:单元测试)。AMD Strix Halo,64 GB,Radeon 8060S,Windows,Vulkan。

构建 权重来自 文件 加载时占用的池 任务 1-8 任务 9-20 任务 21-32 任务 9-20 的 token 数 任务 9-20 的平均解码速度 备注
Swift 3-bit Flash-Next 的 Swift 1.5 微调版本 58.2 GiB ~41 GiB(30 专用 + 10-12 共享) 8/8 11/12 9/12 98,937 28 tok/s 放得下时推荐
Swift 2.5-bit (32 GB) Flash-Next 的 Swift 1.5 微调版本 53 GiB 可放进 32 GB GPU(GPU 权重约 26 GiB) 7/8 10/12 10/12 85,714 33 tok/s 用于 32 GB 显卡
ISTA 3-bit 基础模型(未微调)的 ISTA-DASLab GSQ-RCO 58.2 GiB ~41 GiB(30 专用 + 11 共享) 8/8 11/12 9/12 100,341 28 tok/s 基础模型权重
ISTA + Darwin-R3 (experimental) 嫁接了 496 个 Darwin-180B-RSI-R3 张量的 ISTA 构建 58.7 GiB ~41 GiB(30 专用 + 11 共享) 8/8 11/12 10/12 89,952 30 tok/s 实验性,没有证明有优势

在一台机器上的单次运行:相差一个任务属于噪声范围;token 数和解码速度是实测的。Swift 3 比特构建还通过了在 -ub 512 下的 131072 个 token 的上下文测试(在 18k 到 118k 个 token 处检索五个预埋的事实),以及一个淘汰赛式的编程测试(约 99/100)。 深度审计类测试、德语探测和困惑度见各个模型卡。

运行时说明

  • 共享 MTP 头只能在 Unsloth 分支中加载(danielhanchen/llama.cpp,分支 qwen4exp/mtp,提交 6fcaa16;我们使用 MSVC 和 Vulkan SDK 1.4.357 为 Vulkan 构建了它)。官方 llama.cpp b11259/b11294 在没有推测解码的情况下运行这些文件(约 24 tok/s),因为那里没有仅含 MTP 的头的加载器。
  • 另一个社区构建(agentionai/llama.cpp)使用它自己的 MTP 草稿文件;这两种头在两个构建之间 不能互换。
  • 请使用 -lm dio --lazy-mode on,使逐层 token 表被懒加载读取。对于 131072 的上下文,请使用 -ub 512(-ub 2048 在那里会耗尽 Vulkan 内存)。
  • 在 Windows 上,一旦常驻权重超过专用 GPU 内存,ROCm 的速度就会骤降到约 6 tok/s;上述 Vulkan 数据包含了溢出到共享内存的情况。

工具与数据(本仓库)

  • tools/prune_raw_gguf.py:在字节级别裁剪任何 qwen4exp GGUF 的专家(适用于每种张量类型,包括 gguf-py 不认识的类型)。
  • tools/graft_tensors.py:把 GGUF 中指定名称的张量替换为通过 HTTP range 请求从另一个 GGUF 获取的张量(用于 Darwin-R3 构建);tools/hf_gguf_header.py 读取远程 GGUF 的张量表。
  • keep_lists/:K320、K336、K352、K368、K384、K416、K448 的每层保留专家(原始编号);嵌套(每个较小的集合都包含在下一个之中)。按路由质量排序(在德语或英语中是热点),基于德语 + 英语的校准混合文本。

坦率的局限

不经再训练的专家剪枝会损失质量;我们只在德语/英语文本、困惑度和自己的编程任务上检查过。以上全部是单次运行。这里没有任何东西是一个基准测试套件。

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including webmp3/Sakura-Qwen3.8-Flash-Next-Overview