tachyone-multi (System One decision engine)
Status: released (
v0.4.0), revision 2026-09-29. Trained on a single RTX 3060 12GB and published as LoRA adapters (munod/tachyone-en,munod/tachyone-multi); measured numbers below come frombenchmarks/report.md.This revision is B-11 + B-12 (ADR-0014, ADR-0015): every label in all three primitives is derived from the text it accompanies β
noulfrom its phrase bank (requestβ 1,neutral/empty β 0),scorefrom the tone's level (empty β middle),choicefrom the option the state names (empty β the catch-allother). All datasets were regenerated and the label audit published with the evaluation reports 0 contradictory rows. These numbers are not comparable with pre-B-11/pre-B-12 measurements: the old labels contradicted 121 of 241 request-toned English rows, left everynoullabel ines/de/nlat 0, and gave 7.8% ofscorerows a "near-tie" the text never showed.Provenance, stated plainly: the English adapter carries the B-11 weights and the multilingual adapter the B-12 retrain β each is the best measured checkpoint of its recipe (training on the corrected labels makes
noul+scoretrivial and costschoice; an identical-recipe control landed 13 points lower, L-006).
Model details
- Developed by: The Tachyone Authors.
- Model type: non-autoregressive encoder with three task distributions (
noul,choice,score), answering typed questions about a state in one forward pass. - Trunk: ModernBERT-large (English) and mmBERT-base (100+ languages); see ADR-0007.
- Adapters:
munod/tachyone-en,munod/tachyone-multi(LoRA; load base + adapter). - License: Apache-2.0.
- Repository: https://github.com/munod/tachyone
Uses
Tachyone answers atomic choice / score / noul questions about a state and returns typed values
with probabilities and confidence. It speaks the TypeSafe Jev /v1/systemone wire protocol as
a drop-in and runs locally/offline with no API key. Compose several atomic answers in code
rather than asking one broad question.
Out of scope: free-form text generation, multi-step reasoning, and any decision requiring extended deliberation β decompose those into atomic questions and combine results in code.
Bias, risks, and limitations
- Probabilities are only meaningful after calibration; the shipped temperature must be
applied (see
docs/training.md). - Synthetic training data can inherit generator biases; public probes are evaluation-only.
- Confidence is a property of the distribution, not a guarantee of correctness.
Training
Deterministic synthetic JSONL (training/generate_data.py) supervised with an RLCD
proper-scoring objective (training/finetune_rlcd.py), then temperature-calibrated on a held-out
split (training/fit_calibration.py). Configs and seed live under training/configs/.
Evaluation
Reported by training/evaluate.py and rendered by benchmarks/report.py (accuracy, ECE, p50/p95
latency per primitive and language).
Full-scale run (single RTX 3060 12GB): 9,000 English / 18,000 multilingual train / 1,500 eval
deterministic synthetic records (fully localized per language, a learnable other team with rich
descriptions, per-record RNG, one-in-six distractor clauses), LoRA (r=16 English, r=64 multilingual)
plus a dedicated low-rank choice head (r=32, near-identity init); 4 epochs for English and
8 for multilingual, batch 16, bf16 + gradient checkpointing.
| Checkpoint | Accuracy | ECE (calibrated) | p50 (ms) |
|---|---|---|---|
| English (ModernBERT-large + LoRA r=16 + choice head) | 0.972 | 0.020 | 23.1 |
| Multilingual (mmBERT-base + LoRA r=64 + choice head) | 0.743 | 0.089 | 16.0 |
Per primitive (English): choice 0.946, noul 0.992, score 0.978; (multilingual): choice
0.468, noul 0.892, score 0.870. Label audit: every noul row is judged against its own
text β 0 contradictory in both eval sets (positive rates 0.482 / 0.486), per language in
benchmarks/report.md.
choice is the weak primitive on the multilingual side (0.468) and it is a training trade, not
a label problem: choice labels never changed, and retraining on the corrected labels makes
noul (1.000) and score (0.998) trivial β the same tone detector β while the shared trunk
starves choice (an identical-recipe English control landed at 0.841, the five-domain retrain's
choice collapsed to 0.303). That is why the published English adapter keeps its B-11 weights
(provenance stated in the header). One of six languages meets ECE β€ 0.05 (es 0.045); pt 0.062,
fr 0.101, de 0.118, nl 0.192 and it 0.197 remain above target (NFR-C06), and multilingual
choice/noul ECE (0.099 / 0.108) is declared with them. The CUDA-graph fast path
(TACHYONE_FAST=1) gives a 2.65Γ p50 speedup (8.97 β 3.39 ms) with 0 top-label flips.
Robustness (B-4). On a noisy view (one surface edit β typo/accents/casing β applied to 15% of states) English drops only 0.972 β 0.969 and multilingual 0.743 β 0.744, so the released adapters are robust to this noise model.
Full tables and environment are in
benchmarks/report.md.
Citation
@misc{tachyone2026,
title = {tachyone: a local-first System One decision engine},
author = {The tachyone Authors},
year = {2026},
howpublished = {\url{https://github.com/munod/tachyone}}
}