bob-int
The weights and every file in this repository are licensed Apache-2.0. The bob name and marks are not part of that grant.
Model: bob 2-ck_13000 (bob-talk)
A small integer language model, designed to be checked not used.
Try it: https://bob-talk.org ; the model in this repository is the model answering that page, verifiable by digest (below).
Bob exists because I asked the question: "If fixed point buys exactness, can exact multithreading/parallelism then provide speed?" Bob has demonstrated the ability to resume his own training, cross-ISA and byte-exact reproducible. I do not know if Bob's training model can do this on every possible combination- and make no claims to his ability outside of what I have already demonstrated.
What is in this repository
ck_13000.weights.bin the final weights, training step 13,000
sha256 53ae7609f22e580ac1d67ad1995dfc2d87c0ebcd5f60f3ab7136345f502c2890
ck_13000.opt.bin the optimizer state at the same step, so the
trajectory can be resumed, not only replayed
registration/exponents.json the registered fixed-point exponents and bit
widths for every activation and gradient site
registration/width-sites.json the accumulator width registrations
registration/init/ the initial weights and ordering- the
trajectory is replayable from step 0
SHA256SUMS digests for every file above
The property these files carry
Every number in training and inference is fixed-point integer arithmetic under a registered contract: each site has a registered exponent and width, and arithmetic that would leave its registration refuses- it stops with a named site and an exit code rather than saturating or wrapping silently.
Because the arithmetic is integer, results do not depend on floating-point summation order. The digest of the reply tokens for a given (prompt, seed, effort) is the same on every conforming build. Four pre-published examples with their digests are on https://bob-talk.org and embedded in its page source.
Two portability results are in the project record, each stated no wider than its measurement:
bob 2's own training (the weights in this repository) was saved mid-run and resumed on different rented machines across two GPU generations of one vendor, with checkpoint digests matching at the joints. All of bob 2's GPU lineages are one vendor; no cross-vendor claim is made for this trajectory.
A separate 80-step synthetic trial (2026-08-17, not these weights)
ran the same integer training on an Apple M4 via Metal (arm64, macOS)
and a Pascal-era NVIDIA GTX 1070 via Vulkan (x86_64, Linux)-> produced
byte-identical weights AND optimizer state- sha256
cddab458c1e2e8b9… / 436114216b977be2…- hashed independently on
each machine. The two devices also made byte-identical refusal
decisions: the enforcement machinery itself is deterministic across
vendors. The trial is synthetic and unregistered; it demonstrates the
arithmetic contract's portability, not this model's history.
I know of no shipped system combining an end-to-end integer training loop with integer optimizer state, overflow semantics that refuse with a named site and typed exit, and demonstrated mid-run save-and-resume with matching digests- resting on a documented 2026-08-17 survey of eleven named systems: unoccupied as of the search date, not impossible, not permanent. Please direct me toward prior art if available.
Verify this is the model behind bob-talk.org
sha256sum ck_13000.weights.bin
# 53ae7609f22e580ac1d67ad1995dfc2d87c0ebcd5f60f3ab7136345f502c2890
That digest is compiled into the public server, which refuses to start
if its embedded weights do not hash to it. Ask the live page
(POST /ask, {"prompt": "Alice was", "seed": "7"}) and compare the
reply digest against the published table on the page.
Format
Weights are raw little-endian integer tensors in the registered
layout; registration/exponents.json gives each site's fixed-point
interpretation. This is not a transformers/safetensors checkpoint and
will not load with from_pretrained. Geometry: d_model 256, 5 layers,
8 heads, head_dim 32, d_mlp 1024, context 512 (position table 8192),
vocab 4097; parameter count 5,129,216 (registration/width-sites.json
geometry; ck_13000.weights.bin is 20,516,864 bytes = 5,129,216 x 4).
What is NOT claimed
bob 2 is deterministic in the sense stated above. He is not thereby accurate, clever, or safe, and no such claim is made. He is small (5.1M parameters), his prose is what a 5.1M-parameter model's prose is, and the property I claim to be interesting is the arithmetic, not the model's vocabulary.
Provenance
Trained August of 2026 by f32. Training step 13,000; two legs; the enforcement regime is part of the record (one leg ended in a registered disagreement event on a named site; the refusal machinery was demonstrated live, then the registration corrected and the run completed).
Erratum
2026-09-02: an earlier version of this card stated d_model 512, 8 layers, vocab 8192 and 33,608,704 parameters. Those figures describe a different configuration, not the files in this repository. The card now states the shipped registration; the digest bob-talk.org serves is the truth.
- Downloads last month
- -