YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
binllm
A from-scratch C++20 implementation of a strict 1-bit (binary, {-1,+1}) weight large language model in the BitNet 1.0 regime, with packed weights, SIMD GEMMs, an STE-based training loop, and an inference mode. Modular, parameterized presets from 125M to 70B.
Status — read this first
This tree is a draft that has NOT been compiled or run. It was written against the BitNet 1.0 paper and reference kernel practice, but no compiler has seen it and no test has executed. Treat every claim below as the design, not the observed. The first real check is:
cmake -B build && cmake --build build -j && ctest --test-dir build
The two most valuable tests are test_pack (pins the bit-level packing contract) and
test_kernels (checks every SIMD kernel against the scalar reference, plus a BitLinear
learning smoke that proves the STE wiring carries gradients).
What it implements
| Piece | Where | Paper basis |
|---|---|---|
BitLinear: y = W̃·Quant(LN(x))·(βγ/Qb), W̃ = Sign(W−α), α = mean(W), β = mean|W| |
nn/bitlinear.* |
BitNet 1.0 §2.1, Eqs. 1–12 (arXiv 2310.11453) |
| SubLN — LayerNorm inside each projection, before 8-bit absmax activation quantization | BitLinear::forward |
Eqs. 8–10, Table 4 (20.34 PPL vs Pre-LN 22.11 at 700M) |
| Identity STE through Sign and Clip; fp32 latent weights re-binarized per step | nn/bitlinear.* |
§2.2 (no hard-tanh mask — not BinaryConnect) |
W1A8 GEMM: int8 activations × 1-bit weights via vpmaddubsw/vpsignb |
kernels/gemm_* |
W1A8 is the quality-preserving path (Table 3: 6.7B PPL 17.07 vs FP16 15.19) |
| W1A1 XNOR+popcount, exact 64-bit padding correction | kernels/gemm_* |
included for experiments; binarized activations cost ~5+ pts accuracy |
| Runtime ISA dispatch: AVX2, AVX-512 (incl. VPOPCNTDQ), NEON, scalar | kernels/gemm_dispatch.cpp |
per-ISA TUs so an AVX-512 build can't poison an old CPU |
| SwiGLU MLP, GQA attention, RoPE (+inverse), KV cache | nn/*, model/* |
LLaMA-style head; BitNet Table 5 shapes |
| Adam(0.9, 0.98), poly decay, warmup 750, no grad clipping by default | optim/* |
Table 8; peak LR ~10× fp16 recipe (§3.4) |
| Pile-style uint16 token shards (mmap), micro-batching + grad accumulation | data/*, train/* |
SentencePiece 16K vocab per paper |
| Top-k / top-p / temperature sampling, per-token activation quantization | infer/* |
per-token inference mode (§2.1) |
fp32 latent checkpoints + deployable W1 artifact (packed weights + β + LN affine) |
model/transformer.* |
what an inference runtime loads |
Algorithm contract (the details that matter)
- Packing:
bit = (w < 0), LSB-first, 64 weights peruint64_t; rows padded to a word boundary. Padded bits are 0 (+1) and the GEMMs subtract their contribution exactly:dot = 2·popcount(~(A^W)) − 2·64nwords + K. - Scalar convention: per-output-channel α and β (XNOR-Net / bitnet.cpp convention). BitNet Eq. 3 defines β as one per-layer scalar; setting all channels equal recovers it.
- Activation quantization: symmetric 8-bit absmax,
Qb = 128, clipped to ±127 (the W1A8 kernel usesvpsignb, and|−128|is not representable in int8 — the −128 bucket is intentionally unused). - STE placement: gradients flow unmodified through Sign and the quantizer; the backward of BitLinear is the ordinary linear-layer backward against the quantized input. The per-channel scale multiplies the upstream gradient before the GEMMs.
- No biases on any projection (LLaMA lineage; the paper specifies none).
- Weight tying is on by default (
tie_embeddings): embedding and head share storage.
Build
git clone https://huggingface.co/dosier/binllm
cd binllm
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j
ctest --test-dir build # test_pack, test_kernels
./build/train_tiny runs/tiny # ~1M-param smoke; writes .ckpt + .w1
./build/generate runs/tiny/tiny.w1 1 2 3 --n 16
Requires CMake ≥ 3.20, a C++20 compiler (GCC/Clang), POSIX mmap (Linux/macOS;
Windows not supported yet). No third-party runtime dependencies.
Known deviations and caveats
- Not compiled/validated yet (see Status). The first
cmake --buildis expected to surface minor issues; the design contract above is what the tests pin. - fp32
gemm_fp32routestransB=trueto the scalar loop (only used off the hot path). - The forward pass is single-threaded per call; parallelize at the batch level.
- The trainer treats the batch as one flat causal sequence (no cross-sample masking), the standard fast-pretraining simplification.
Layout
include/binllm/{config,core,data,kernels,nn,model,optim,train,infer}
src/{core,kernels,nn,model,train,infer}
tests/, examples/
License
MIT.