system-one-270m

A System One model: it takes a piece of state plus a typed question with a caller-supplied option set, and returns one option with a calibrated probability. It never generates free text.

Built as an open reproduction of the idea behind TypeSafe's Jev (closed) and convaiinnovations/laya (ModernBERT-large), on a 270M decoder rather than an encoder.

How it answers

Options are rendered as letters in the prompt. After Gemma's <start_of_turn>model\n a bare letter is exactly one token, so a single forward pass yields logits over the whole answer space at once. The head gathers the first n_options letter ids, masks the rest to -inf, and softmaxes over that slice.

  • Type safety โ€” an option outside the caller's set is not representable.
  • Calibration โ€” trained against soft targets with the log score, a strictly proper scoring rule, so honest probabilities maximise the objective.
  • Flat latency in option count โ€” options are input, not N generations.
import torch, torch.nn.functional as F
from transformers import AutoModelForCausalLM, AutoTokenizer

LETTERS = "ABCDEFGHIJKLMNOPQRSTUVWXYZ"
tok = AutoTokenizer.from_pretrained("kaivoss/system-one-270m")
model = AutoModelForCausalLM.from_pretrained("kaivoss/system-one-270m").eval()

def decide(state, question, options):
    body = "\n".join(f"{LETTERS[i]}. {o['label']} - {o['description']}"
                      for i, o in enumerate(options))
    prompt = (f"<state>\n{state}\n</state>\n\nQuestion: {question}\n"
              f"Options:\n{body}\n\nAnswer with one letter.\nAnswer:")
    text = tok.apply_chat_template([{"role": "user", "content": prompt}],
                                   tokenize=False, add_generation_prompt=True)
    ids = torch.tensor([tok.encode(text, add_special_tokens=False)])
    with torch.no_grad():
        logits = model(input_ids=ids, logits_to_keep=1).logits[:, -1, :]
    letter_ids = torch.tensor([tok.encode(LETTERS[i], add_special_tokens=False)[0]
                               for i in range(len(options))])
    probs = F.softmax(logits.index_select(1, letter_ids), dim=-1)[0]
    best = int(probs.argmax())
    return options[best]["label"], float(probs[best])

print(decide(
    "My card was charged twice for the same month and support hasn't replied.",
    "Which team should handle this?",
    [{"label": "billing",   "description": "Payment or subscription issues"},
     {"label": "technical", "description": "Bugs or integration problems"},
     {"label": "sales",     "description": "Pricing or account questions"}]))

Results

2,493 held-out questions, split by state so no state appears in both train and eval.

Metric Base (gemma-3-270m-it) This model
Accuracy 0.4204 0.6574
Brier 0.7182 0.4110
ECE 0.3035 0.1311
ECE (temperature-scaled, T=2.0) 0.1062 0.0374
Random baseline 0.3471 0.3471

Calibration matters more than accuracy for this model class. A model right 66% of the time that says 66% is usable; one that says 99% is not. The ECE drop is what the proper-scoring-rule objective bought.

Training data

25,002 synthetic typed decisions over 7,537 states, generated with openai/gpt-oss-20b via OpenRouter at ~$0.000275/question. 28 domains; 49.3% choice, 33.4% noul (yes/no), 17.3% score (ordinal).

Soft targets come from an ensemble of permuted reads: gpt-oss reasons before answering and cannot be told not to, which makes any single response's logprobs effectively one-hot. The uncertainty lives in the variance across reads, and shuffling option order each time also cancels letter-position bias.

Training: 3 epochs, bs 32, lr 5e-5 cosine, ~24 min on one H100 SXM.

Limitations

  • 84.6% of training rows are unambiguous, so graded uncertainty is under-represented and calibration on genuinely contested inputs is the weakest part of the model.
  • Accuracy is below Laya's reported 0.766 (on its own benchmark, which is not this eval set โ€” the two numbers are not directly comparable).
  • Option sets above ~14 are barely represented (2.6% of rows at 9โ€“14, 0.04% above 15); the letter scheme caps at 26.
  • Labels are model-generated, not human, so the ceiling is gpt-oss-20b's own judgement, and any systematic bias it has is baked in.
  • State is not treated as hostile. Text inside the state can steer the answer; this is an injection surface, exactly as TypeSafe documents for Jev.
  • English only in practice, though nothing in the design is language-specific.
Downloads last month
69
Safetensors
Model size
0.3B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for kaivoss/system-one-270m

Finetuned
(419)
this model