YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

binllm

A from-scratch C++20 implementation of a strict 1-bit (binary, {-1,+1}) weight large language model in the BitNet 1.0 regime, with packed weights, SIMD GEMMs, an STE-based training loop, and an inference mode. Modular, parameterized presets from 125M to 70B.

Status — read this first

This tree is a draft that has NOT been compiled or run. It was written against the BitNet 1.0 paper and reference kernel practice, but no compiler has seen it and no test has executed. Treat every claim below as the design, not the observed. The first real check is:

cmake -B build && cmake --build build -j && ctest --test-dir build

The two most valuable tests are test_pack (pins the bit-level packing contract) and test_kernels (checks every SIMD kernel against the scalar reference, plus a BitLinear learning smoke that proves the STE wiring carries gradients).

What it implements

Piece Where Paper basis
BitLinear: y = W̃·Quant(LN(x))·(βγ/Qb), W̃ = Sign(W−α), α = mean(W), β = mean|W| nn/bitlinear.* BitNet 1.0 §2.1, Eqs. 1–12 (arXiv 2310.11453)
SubLN — LayerNorm inside each projection, before 8-bit absmax activation quantization BitLinear::forward Eqs. 8–10, Table 4 (20.34 PPL vs Pre-LN 22.11 at 700M)
Identity STE through Sign and Clip; fp32 latent weights re-binarized per step nn/bitlinear.* §2.2 (no hard-tanh mask — not BinaryConnect)
W1A8 GEMM: int8 activations × 1-bit weights via vpmaddubsw/vpsignb kernels/gemm_* W1A8 is the quality-preserving path (Table 3: 6.7B PPL 17.07 vs FP16 15.19)
W1A1 XNOR+popcount, exact 64-bit padding correction kernels/gemm_* included for experiments; binarized activations cost ~5+ pts accuracy
Runtime ISA dispatch: AVX2, AVX-512 (incl. VPOPCNTDQ), NEON, scalar kernels/gemm_dispatch.cpp per-ISA TUs so an AVX-512 build can't poison an old CPU
SwiGLU MLP, GQA attention, RoPE (+inverse), KV cache nn/*, model/* LLaMA-style head; BitNet Table 5 shapes
Adam(0.9, 0.98), poly decay, warmup 750, no grad clipping by default optim/* Table 8; peak LR ~10× fp16 recipe (§3.4)
Pile-style uint16 token shards (mmap), micro-batching + grad accumulation data/*, train/* SentencePiece 16K vocab per paper
Top-k / top-p / temperature sampling, per-token activation quantization infer/* per-token inference mode (§2.1)
fp32 latent checkpoints + deployable W1 artifact (packed weights + β + LN affine) model/transformer.* what an inference runtime loads

Algorithm contract (the details that matter)

  • Packing: bit = (w < 0), LSB-first, 64 weights per uint64_t; rows padded to a word boundary. Padded bits are 0 (+1) and the GEMMs subtract their contribution exactly: dot = 2·popcount(~(A^W)) − 2·64nwords + K.
  • Scalar convention: per-output-channel α and β (XNOR-Net / bitnet.cpp convention). BitNet Eq. 3 defines β as one per-layer scalar; setting all channels equal recovers it.
  • Activation quantization: symmetric 8-bit absmax, Qb = 128, clipped to ±127 (the W1A8 kernel uses vpsignb, and |−128| is not representable in int8 — the −128 bucket is intentionally unused).
  • STE placement: gradients flow unmodified through Sign and the quantizer; the backward of BitLinear is the ordinary linear-layer backward against the quantized input. The per-channel scale multiplies the upstream gradient before the GEMMs.
  • No biases on any projection (LLaMA lineage; the paper specifies none).
  • Weight tying is on by default (tie_embeddings): embedding and head share storage.

Build

git clone https://huggingface.co/dosier/binllm
cd binllm
cmake -B build -DCMAKE_BUILD_TYPE=Release
cmake --build build -j
ctest --test-dir build            # test_pack, test_kernels
./build/train_tiny runs/tiny      # ~1M-param smoke; writes .ckpt + .w1
./build/generate runs/tiny/tiny.w1 1 2 3 --n 16

Requires CMake ≥ 3.20, a C++20 compiler (GCC/Clang), POSIX mmap (Linux/macOS; Windows not supported yet). No third-party runtime dependencies.

Known deviations and caveats

  • Not compiled/validated yet (see Status). The first cmake --build is expected to surface minor issues; the design contract above is what the tests pin.
  • fp32 gemm_fp32 routes transB=true to the scalar loop (only used off the hot path).
  • The forward pass is single-threaded per call; parallelize at the batch level.
  • The trainer treats the batch as one flat causal sequence (no cross-sample masking), the standard fast-pretraining simplification.

Layout

include/binllm/{config,core,data,kernels,nn,model,optim,train,infer}
src/{core,kernels,nn,model,train,infer}
tests/, examples/

License

MIT.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support