tachyone-multi (System One decision engine)

Status: released (v0.4.0), revision 2026-09-29. Trained on a single RTX 3060 12GB and published as LoRA adapters (munod/tachyone-en, munod/tachyone-multi); measured numbers below come from benchmarks/report.md.

This revision is B-11 + B-12 (ADR-0014, ADR-0015): every label in all three primitives is derived from the text it accompanies β€” noul from its phrase bank (request β†’ 1, neutral/empty β†’ 0), score from the tone's level (empty β†’ middle), choice from the option the state names (empty β†’ the catch-all other). All datasets were regenerated and the label audit published with the evaluation reports 0 contradictory rows. These numbers are not comparable with pre-B-11/pre-B-12 measurements: the old labels contradicted 121 of 241 request-toned English rows, left every noul label in es/de/nl at 0, and gave 7.8% of score rows a "near-tie" the text never showed.

Provenance, stated plainly: the English adapter carries the B-11 weights and the multilingual adapter the B-12 retrain β€” each is the best measured checkpoint of its recipe (training on the corrected labels makes noul+score trivial and costs choice; an identical-recipe control landed 13 points lower, L-006).

Model details

  • Developed by: The Tachyone Authors.
  • Model type: non-autoregressive encoder with three task distributions (noul, choice, score), answering typed questions about a state in one forward pass.
  • Trunk: ModernBERT-large (English) and mmBERT-base (100+ languages); see ADR-0007.
  • Adapters: munod/tachyone-en, munod/tachyone-multi (LoRA; load base + adapter).
  • License: Apache-2.0.
  • Repository: https://github.com/munod/tachyone

Uses

Tachyone answers atomic choice / score / noul questions about a state and returns typed values with probabilities and confidence. It speaks the TypeSafe Jev /v1/systemone wire protocol as a drop-in and runs locally/offline with no API key. Compose several atomic answers in code rather than asking one broad question.

Out of scope: free-form text generation, multi-step reasoning, and any decision requiring extended deliberation β€” decompose those into atomic questions and combine results in code.

Bias, risks, and limitations

  • Probabilities are only meaningful after calibration; the shipped temperature must be applied (see docs/training.md).
  • Synthetic training data can inherit generator biases; public probes are evaluation-only.
  • Confidence is a property of the distribution, not a guarantee of correctness.

Training

Deterministic synthetic JSONL (training/generate_data.py) supervised with an RLCD proper-scoring objective (training/finetune_rlcd.py), then temperature-calibrated on a held-out split (training/fit_calibration.py). Configs and seed live under training/configs/.

Evaluation

Reported by training/evaluate.py and rendered by benchmarks/report.py (accuracy, ECE, p50/p95 latency per primitive and language).

Full-scale run (single RTX 3060 12GB): 9,000 English / 18,000 multilingual train / 1,500 eval deterministic synthetic records (fully localized per language, a learnable other team with rich descriptions, per-record RNG, one-in-six distractor clauses), LoRA (r=16 English, r=64 multilingual) plus a dedicated low-rank choice head (r=32, near-identity init); 4 epochs for English and 8 for multilingual, batch 16, bf16 + gradient checkpointing.

Checkpoint Accuracy ECE (calibrated) p50 (ms)
English (ModernBERT-large + LoRA r=16 + choice head) 0.972 0.020 23.1
Multilingual (mmBERT-base + LoRA r=64 + choice head) 0.743 0.089 16.0

Per primitive (English): choice 0.946, noul 0.992, score 0.978; (multilingual): choice 0.468, noul 0.892, score 0.870. Label audit: every noul row is judged against its own text β€” 0 contradictory in both eval sets (positive rates 0.482 / 0.486), per language in benchmarks/report.md.

choice is the weak primitive on the multilingual side (0.468) and it is a training trade, not a label problem: choice labels never changed, and retraining on the corrected labels makes noul (1.000) and score (0.998) trivial β€” the same tone detector β€” while the shared trunk starves choice (an identical-recipe English control landed at 0.841, the five-domain retrain's choice collapsed to 0.303). That is why the published English adapter keeps its B-11 weights (provenance stated in the header). One of six languages meets ECE ≀ 0.05 (es 0.045); pt 0.062, fr 0.101, de 0.118, nl 0.192 and it 0.197 remain above target (NFR-C06), and multilingual choice/noul ECE (0.099 / 0.108) is declared with them. The CUDA-graph fast path (TACHYONE_FAST=1) gives a 2.65Γ— p50 speedup (8.97 β†’ 3.39 ms) with 0 top-label flips.

Robustness (B-4). On a noisy view (one surface edit β€” typo/accents/casing β€” applied to 15% of states) English drops only 0.972 β†’ 0.969 and multilingual 0.743 β†’ 0.744, so the released adapters are robust to this noise model.

Full tables and environment are in benchmarks/report.md.

Citation

@misc{tachyone2026,
  title        = {tachyone: a local-first System One decision engine},
  author       = {The tachyone Authors},
  year         = {2026},
  howpublished = {\url{https://github.com/munod/tachyone}}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support