system-one-270m
A System One model: it takes a piece of state plus a typed question with a caller-supplied option set, and returns one option with a calibrated probability. It never generates free text.
Built as an open reproduction of the idea behind TypeSafe's
Jev (closed)
and convaiinnovations/laya
(ModernBERT-large), on a 270M decoder rather than an encoder.
How it answers
Options are rendered as letters in the prompt. After Gemma's
<start_of_turn>model\n a bare letter is exactly one token, so a single forward
pass yields logits over the whole answer space at once. The head gathers the
first n_options letter ids, masks the rest to -inf, and softmaxes over that
slice.
- Type safety โ an option outside the caller's set is not representable.
- Calibration โ trained against soft targets with the log score, a strictly proper scoring rule, so honest probabilities maximise the objective.
- Flat latency in option count โ options are input, not N generations.
import torch, torch.nn.functional as F
from transformers import AutoModelForCausalLM, AutoTokenizer
LETTERS = "ABCDEFGHIJKLMNOPQRSTUVWXYZ"
tok = AutoTokenizer.from_pretrained("kaivoss/system-one-270m")
model = AutoModelForCausalLM.from_pretrained("kaivoss/system-one-270m").eval()
def decide(state, question, options):
body = "\n".join(f"{LETTERS[i]}. {o['label']} - {o['description']}"
for i, o in enumerate(options))
prompt = (f"<state>\n{state}\n</state>\n\nQuestion: {question}\n"
f"Options:\n{body}\n\nAnswer with one letter.\nAnswer:")
text = tok.apply_chat_template([{"role": "user", "content": prompt}],
tokenize=False, add_generation_prompt=True)
ids = torch.tensor([tok.encode(text, add_special_tokens=False)])
with torch.no_grad():
logits = model(input_ids=ids, logits_to_keep=1).logits[:, -1, :]
letter_ids = torch.tensor([tok.encode(LETTERS[i], add_special_tokens=False)[0]
for i in range(len(options))])
probs = F.softmax(logits.index_select(1, letter_ids), dim=-1)[0]
best = int(probs.argmax())
return options[best]["label"], float(probs[best])
print(decide(
"My card was charged twice for the same month and support hasn't replied.",
"Which team should handle this?",
[{"label": "billing", "description": "Payment or subscription issues"},
{"label": "technical", "description": "Bugs or integration problems"},
{"label": "sales", "description": "Pricing or account questions"}]))
Results
2,493 held-out questions, split by state so no state appears in both train and eval.
| Metric | Base (gemma-3-270m-it) |
This model |
|---|---|---|
| Accuracy | 0.4204 | 0.6574 |
| Brier | 0.7182 | 0.4110 |
| ECE | 0.3035 | 0.1311 |
| ECE (temperature-scaled, T=2.0) | 0.1062 | 0.0374 |
| Random baseline | 0.3471 | 0.3471 |
Calibration matters more than accuracy for this model class. A model right 66% of the time that says 66% is usable; one that says 99% is not. The ECE drop is what the proper-scoring-rule objective bought.
Training data
25,002 synthetic typed decisions over 7,537 states, generated with
openai/gpt-oss-20b via OpenRouter at ~$0.000275/question. 28 domains;
49.3% choice, 33.4% noul (yes/no), 17.3% score (ordinal).
Soft targets come from an ensemble of permuted reads: gpt-oss reasons before answering and cannot be told not to, which makes any single response's logprobs effectively one-hot. The uncertainty lives in the variance across reads, and shuffling option order each time also cancels letter-position bias.
Training: 3 epochs, bs 32, lr 5e-5 cosine, ~24 min on one H100 SXM.
Limitations
- 84.6% of training rows are unambiguous, so graded uncertainty is under-represented and calibration on genuinely contested inputs is the weakest part of the model.
- Accuracy is below Laya's reported 0.766 (on its own benchmark, which is not this eval set โ the two numbers are not directly comparable).
- Option sets above ~14 are barely represented (2.6% of rows at 9โ14, 0.04% above 15); the letter scheme caps at 26.
- Labels are model-generated, not human, so the ceiling is gpt-oss-20b's own judgement, and any systematic bias it has is baked in.
- State is not treated as hostile. Text inside the state can steer the answer; this is an injection surface, exactly as TypeSafe documents for Jev.
- English only in practice, though nothing in the design is language-specific.
- Downloads last month
- 69