BitNet b1.58 2B4T — TQ2_0 GGUF (mainline llama.cpp)

TQ2_0 conversion of Microsoft's BitNet b1.58 2B4T (MIT), converted on mainline llama.cpp with synapticode-ai/llama.cpp (two small converter patches; conversion walkthrough in docs/tq2_0-anatomy.md).

bf16-equivalent quality at 2.06 bits per weight.

File

File Format Size sha256
bitnet-2b4t-tq2_0.gguf TQ2_0 (2.0625 bpw) 1.2 GB 9f8e1097502528a0d80d885c603ea7ee3e4d214a6685356e39baaf697c02cbb6

Conversion is deterministic: a fresh bf16 → TQ2_0 conversion on the pinned tree reproduces this file byte-for-byte.

Quality (perplexity, measured)

llama-perplexity, WikiText-2 raw test set (corpus sha256 bbf94c53a05abe9e…), CPU (-ngl 0), substrate = synapticode-ai/llama.cpp v0.1.0. Same corpus, commands, and binary for both columns.

n_ctx bf16 reference TQ2_0 (this file) relative Δ
512 82.09 ± 0.76 82.21 ± 0.77 +0.15%
2048 77.16 ± 0.71 77.31 ± 0.71 +0.19%

2B4T is QAT-trained natively ternary; TQ2_0 is its native alphabet. Absolute values reflect an instruction-tuned model scored on raw text and are comparable only within this methodology.

Throughput (measured, hardware disclosed)

Mac mini M4 Pro, CPU-only (-ngl 0; TQ2_0 has no Metal path): 237–279 tok/s prompt eval, 89–112 tok/s generation.

Run

./build/bin/llama-cli -m bitnet-2b4t-tq2_0.gguf -ngl 0 -p "hello"

Source model © Microsoft, MIT. Conversion: Synapticode · code@synapticode.ai

Downloads last month
220
GGUF
Model size
2B params
Architecture
bitnet
Hardware compatibility
Log In to add your hardware

2-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Synapticode/bitnet-b1.58-2B-4T-tq2_0-gguf

Quantized
(8)
this model