Sakura — Qwen3.8-Flash-Next builds: overview
All of our Qwen3.8-Flash-Next GGUF builds in one place: what each one is, how it differs, what we measured, and how to run it. No weights are in this repository.
Qwen3.8-Flash-Next (qwen4exp, 512 experts, 10 active, a 51B-parameter per-layer token table that stays on disk) is large; all builds below keep 352 of the 512 experts per layer (chosen on German + English routing) so that they run on a single 64 GB AMD Strix Halo.
They are independent community builds, not official Qwen releases. Part of the Sakura Micro collection.
Which one should I take?
| Your situation | Take |
|---|---|
| 32 GB GPU | Swift 2.5-bit (about 26 GiB of GPU weights) |
| 64 GB unified memory (Strix Halo class), best coding result and shortest thinking | Swift 3-bit |
| You want the unmodified base-model weights (no fine-tune) | ISTA 3-bit |
| You want to try Darwin-R3 weights on the same layout | ISTA + Darwin-R3 (experimental: in our tests no difference to the ISTA build) |
What differs between the builds
Four independent axes: where the weights come from (Swift fine-tune, ISTA base-model quantization, ISTA with Darwin-R3 tensors), bits (3-bit IQ3_XXS or 2.5-bit IQ2_XS), how many experts (all 352 here) and which experts (the same German + English selection, lists below).
Naming scheme: Sakura-Qwen3.8-Flash-Next-<source>-<bits>-<experts>E[-<size class>]-GGUF.
Measured results (same runtime, tasks and seeds for every build)
Runtime: Unsloth fork qwen4exp/mtp (commit 6fcaa16, Vulkan, built by us), shared MTP head mtp-Qwen3.8-Flash-Next-shared-Q4_K_M.gguf (n-max 2), -b 2048 -ub 2048, 32k context, KV q8_0, thinking budget 10000, max 12288 tokens, temperature 0.6, top-p 0.95, top-k 20, seeds 1000 + task number, one run per task, own frozen task set of 32 LeetCode/AtCoder-style tasks (grader: unit tests). AMD Strix Halo, 64 GB, Radeon 8060S, Windows, Vulkan.
| Build | Weights from | File | Pool at load | tasks 1-8 | tasks 9-20 | tasks 21-32 | tokens 9-20 | avg decode in tasks 9-20 | Note |
|---|---|---|---|---|---|---|---|---|---|
| Swift 3-bit | Swift 1.5 fine-tune of Flash-Next | 58.2 GiB | ~41 GiB (30 dedicated + 10-12 shared) | 8/8 | 11/12 | 9/12 | 98,937 | 28 tok/s | recommended when it fits |
| Swift 2.5-bit (32 GB) | Swift 1.5 fine-tune of Flash-Next | 53 GiB | fits a 32 GB GPU (about 26 GiB of GPU weights) | 7/8 | 10/12 | 10/12 | 85,714 | 33 tok/s | for 32 GB cards |
| ISTA 3-bit | ISTA-DASLab GSQ-RCO of the base model (no fine-tune) | 58.2 GiB | ~41 GiB (30 dedicated + 11 shared) | 8/8 | 11/12 | 9/12 | 100,341 | 28 tok/s | base-model weights |
| ISTA + Darwin-R3 (experimental) | ISTA build with 496 Darwin-180B-RSI-R3 tensors grafted in | 58.7 GiB | ~41 GiB (30 dedicated + 11 shared) | 8/8 | 11/12 | 10/12 | 89,952 | 30 tok/s | experimental, no demonstrated advantage |
Single runs on one machine: a difference of one task is inside the noise; token counts and decode speed are measured. The Swift 3-bit build also passes a 131072-token context test at -ub 512 (retrieval of five planted facts at 18k to 118k tokens) and a knockout-style coding test (about 99/100).
Deep-audit style tests, German probes and perplexity are on the individual model cards.
Runtime notes
- The shared MTP head only loads in the Unsloth fork (
danielhanchen/llama.cpp, branchqwen4exp/mtp, commit 6fcaa16; we built it for Vulkan with MSVC and Vulkan SDK 1.4.357). Official llama.cpp b11259/b11294 run the files without speculation (about 24 tok/s) because the MTP-only head has no loader there. - Another community build (
agentionai/llama.cpp) uses its own MTP draft file; the two heads are not interchangeable between the two builds. - Use
-lm dio --lazy-mode onso that the per-layer token table is read lazily. For 131072 context use-ub 512(-ub 2048runs out of Vulkan memory there). - ROCm on Windows collapses to about 6 tok/s as soon as the resident weights exceed the dedicated GPU memory; the Vulkan numbers above include the spill into shared memory.
Tools and data (this repository)
tools/prune_raw_gguf.py: cuts the experts of anyqwen4expGGUF at the byte level (works for every tensor type, including ones unknown to gguf-py).tools/graft_tensors.py: replaces named tensors of a GGUF by tensors fetched with HTTP range requests from another GGUF (used for the Darwin-R3 build);tools/hf_gguf_header.pyreads a remote GGUF's tensor table.keep_lists/: kept experts per layer (original numbering) for K320, K336, K352, K368, K384, K416, K448; nested (each smaller set is contained in the next). Ranked by routing mass (hot in German or English) on a German + English calibration mix.
Honest limits
Expert pruning without retraining costs quality; we checked it on German/English text, perplexity and our own coding tasks only. Everything above is single runs. Nothing here is a benchmark suite.
中文说明 · 樱花 (Simplified Chinese)
English above. 本节为上文的中文翻译(Sakura = 樱花 yīnghuā);完整的独立中文版见 README_zh.md。
Sakura — Qwen3.8-Flash-Next 构建版本:概览
我们所有 Qwen3.8-Flash-Next GGUF 构建版本集中在一处:每个版本是什么、有何不同、我们测量了什么,以及如何运行。本仓库中没有权重。
Qwen3.8-Flash-Next(qwen4exp,512 个专家,激活 10 个,一个 51B 参数的逐层 token 表留在磁盘上)很大;下面所有构建都在每层 512 个专家中保留 352 个(依据德语 + 英语的路由情况选出),以便在单台 64 GB 的 AMD Strix Halo 上运行。
它们是独立的社区构建版本,不是 Qwen 的官方发布。属于 Sakura Micro 合集。
我该选哪一个?
| 你的情况 | 选择 |
|---|---|
| 32 GB GPU | Swift 2.5 比特(GPU 权重约 26 GiB) |
| 64 GB 统一内存(Strix Halo 级别),追求最佳编程结果和最短的思考 | Swift 3 比特 |
| 你想要未经修改的基础模型权重(未微调) | ISTA 3 比特 |
| 你想在相同布局上试用 Darwin-R3 权重 | ISTA + Darwin-R3(实验性:在我们的测试中与 ISTA 构建没有差别) |
各构建之间有何不同
四个相互独立的维度:权重来自哪里(Swift 微调、ISTA 基础模型量化、带 Darwin-R3 张量的 ISTA)、比特数(3 比特 IQ3_XXS 或 2.5 比特 IQ2_XS)、多少个专家(这里都是 352 个)以及 哪些专家(相同的德语 + 英语选择,列表见下文)。
命名方案:Sakura-Qwen3.8-Flash-Next-<source>-<bits>-<experts>E[-<size class>]-GGUF。
测量结果(每个构建使用相同的运行时、任务和随机种子)
运行时:Unsloth 分支 qwen4exp/mtp(提交 6fcaa16,Vulkan,由我们自己构建),共享 MTP 头 mtp-Qwen3.8-Flash-Next-shared-Q4_K_M.gguf(n-max 2),-b 2048 -ub 2048,32k 上下文,KV q8_0,思考预算 10000,最多 12288 个 token,temperature 0.6,top-p 0.95,top-k 20,随机种子 1000 + 任务编号,每个任务运行一次,使用自己的冻结任务集,共 32 个 LeetCode/AtCoder 风格的任务(评分器:单元测试)。AMD Strix Halo,64 GB,Radeon 8060S,Windows,Vulkan。
| 构建 | 权重来自 | 文件 | 加载时占用的池 | 任务 1-8 | 任务 9-20 | 任务 21-32 | 任务 9-20 的 token 数 | 任务 9-20 的平均解码速度 | 备注 |
|---|---|---|---|---|---|---|---|---|---|
| Swift 3-bit | Flash-Next 的 Swift 1.5 微调版本 | 58.2 GiB | ~41 GiB(30 专用 + 10-12 共享) | 8/8 | 11/12 | 9/12 | 98,937 | 28 tok/s | 放得下时推荐 |
| Swift 2.5-bit (32 GB) | Flash-Next 的 Swift 1.5 微调版本 | 53 GiB | 可放进 32 GB GPU(GPU 权重约 26 GiB) | 7/8 | 10/12 | 10/12 | 85,714 | 33 tok/s | 用于 32 GB 显卡 |
| ISTA 3-bit | 基础模型(未微调)的 ISTA-DASLab GSQ-RCO | 58.2 GiB | ~41 GiB(30 专用 + 11 共享) | 8/8 | 11/12 | 9/12 | 100,341 | 28 tok/s | 基础模型权重 |
| ISTA + Darwin-R3 (experimental) | 嫁接了 496 个 Darwin-180B-RSI-R3 张量的 ISTA 构建 | 58.7 GiB | ~41 GiB(30 专用 + 11 共享) | 8/8 | 11/12 | 10/12 | 89,952 | 30 tok/s | 实验性,没有证明有优势 |
在一台机器上的单次运行:相差一个任务属于噪声范围;token 数和解码速度是实测的。Swift 3 比特构建还通过了在 -ub 512 下的 131072 个 token 的上下文测试(在 18k 到 118k 个 token 处检索五个预埋的事实),以及一个淘汰赛式的编程测试(约 99/100)。
深度审计类测试、德语探测和困惑度见各个模型卡。
运行时说明
- 共享 MTP 头只能在 Unsloth 分支中加载(
danielhanchen/llama.cpp,分支qwen4exp/mtp,提交 6fcaa16;我们使用 MSVC 和 Vulkan SDK 1.4.357 为 Vulkan 构建了它)。官方 llama.cpp b11259/b11294 在没有推测解码的情况下运行这些文件(约 24 tok/s),因为那里没有仅含 MTP 的头的加载器。 - 另一个社区构建(
agentionai/llama.cpp)使用它自己的 MTP 草稿文件;这两种头在两个构建之间 不能互换。 - 请使用
-lm dio --lazy-mode on,使逐层 token 表被懒加载读取。对于 131072 的上下文,请使用-ub 512(-ub 2048在那里会耗尽 Vulkan 内存)。 - 在 Windows 上,一旦常驻权重超过专用 GPU 内存,ROCm 的速度就会骤降到约 6 tok/s;上述 Vulkan 数据包含了溢出到共享内存的情况。
工具与数据(本仓库)
tools/prune_raw_gguf.py:在字节级别裁剪任何qwen4expGGUF 的专家(适用于每种张量类型,包括 gguf-py 不认识的类型)。tools/graft_tensors.py:把 GGUF 中指定名称的张量替换为通过 HTTP range 请求从另一个 GGUF 获取的张量(用于 Darwin-R3 构建);tools/hf_gguf_header.py读取远程 GGUF 的张量表。keep_lists/:K320、K336、K352、K368、K384、K416、K448 的每层保留专家(原始编号);嵌套(每个较小的集合都包含在下一个之中)。按路由质量排序(在德语或英语中是热点),基于德语 + 英语的校准混合文本。
坦率的局限
不经再训练的专家剪枝会损失质量;我们只在德语/英语文本、困惑度和自己的编程任务上检查过。以上全部是单次运行。这里没有任何东西是一个基准测试套件。