Ling-3.0-tiny-sub3bit-PPLp10-GGUF

我的天那,钠是结晶的。
这是我见过最小巧最有用的MoE模型。
所以我想随意的量化它,并为PPL跑分设计。(其实是我想内置到TranslatorMinecraft,因为它是MIT的。)
我会在这个仓库上上传小于3bit且bartowski/Ling-3.0-tiny-calibration-v6.txt验证PPL小于+10%的模型。
校准文件bartowski/Ling-3.0-tiny-imatrix.gguf

00 01 bartowski/Ling-3.0-tiny-bf16.gguf
BPW 2.99 2.98 16
PPL 5.0691±0.03523 5.0480±0.03498 4.6431±0.03271
↑+% 9.1749% 8.7204% 0%
大小 2812.01MiB 2807.11MiB 15065.15MiB
经验 输入相关的不用IQ更强,输出得用IQ。 继左边 空气与Air的混合物
观后感 NAVI回家吧,paiN粘的要死。 呜呜呜我补药开学 网速100Mbps

你懂 01 比 00 小 4.9MiB 是什么感觉吗?换算到 HBM4 价值整整¥1.75。

Ling-3.0-tiny-sub3bit-PPLp10-GGUF

Markdown translation model: DeepSeek V4 Flash 0731.
Oh my god, sodium is crystalline.
This is the smallest and most useful MoE model I have ever seen.
So I decided to casually quantize it and design it for PPL benchmarking. (Actually, I want to embed it into TranslatorMinecraft, because it is MIT-licensed.)
I will upload sub-3-bit models on this repository that have a PPL increase of less than +10%, verified by bartowski/Ling-3.0-tiny-calibration-v6.txt.
Calibration file: bartowski/Ling-3.0-tiny-imatrix.gguf

Item 00 01 bartowski/Ling-3.0-tiny-bf16.gguf
BPW 2.99 2.98 16
PPL 5.0691±0.03523 5.0480±0.03498 4.6431±0.03271
↑+% 9.1749% 8.7204% 0%
Size 2812.01 MiB 2807.11 MiB 15065.15 MiB
Experience For input-related tasks, non-IQ is stronger; for output, use IQ. Continuing from the left A mixture of air and Air
Afterthoughts NAVI, go home; paiN is too sticky/clingy. Boo-hoo, I don't want school to start! Internet speed: 100 Mbps

Can you imagine what it means that 01 is 4.9 MiB less than 00? In HBM4 cost terms, that's a whole $0.25.

Downloads last month
360
GGUF
Model size
8B params
Architecture
bailingmoe3
Hardware compatibility
Log In to add your hardware

We're not able to determine the quantization variants.

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Q1ngMang/Ling-3.0-tiny-sub3bit-PPLp10-GGUF

Quantized
(21)
this model