Darwin-27B-ZTC

A zero-token decision engine from the Darwin family. Darwin-27B-ZTC reads a piece of state and a set of typed questions, and returns a full probability distribution over the options of every question. It does this in one forward pass per question. It generates no tokens.

It takes the POST /v1/systemone request shape: noul (yes or no), choice (one of N labels) and score (an ordered rubric). Every option can carry a written description, and the descriptions are part of the input.

Results

Typed Decisions (general, zero-shot)

LocalLLaMA/typed-decisions, test split: 400 cases, 2,000 decisions. One request per case, with the state and all five questions together, the same request shape as the benchmark README.

Metric Darwin-27B-ZTC
Accuracy ↑ 0.743
KL from gold ↓ 0.204
Brier ↓ 0.097
Question type Decisions Accuracy
noul (yes or no) 600 0.845
choice 600 0.732
score (rubric) 800 0.675

Zero-shot. The model never saw the Typed Decisions train split, its four workflows or its question schemas. All 2,000 decisions were answered, with zero errors.

Scoring follows the benchmark README: accuracy is agreement with the gold label, KL is KL(gold || prediction) averaged over decisions, and Brier is the squared error summed over the options of a decision, averaged over decisions. Our scorer reproduces the README's Uniform reference (KL 0.444, Brier 0.238). We ran the test set twice with the same model and settings. The two runs gave accuracy 0.741 and 0.743. Every number on this card comes from the second run, the one with the saved predictions (noul 507/600, choice 439/600, score 540/800, 1,486/2,000 in total). The two runs differ on 14 of 2,000 decisions (9 better, 5 worse; mean +0.002, 95% case-bootstrap interval -0.002 to +0.006), which is within run-to-run noise: the server batches questions from concurrent requests, so near-tie decisions can flip between runs. An earlier version of this card mixed the per-type numbers of the first run with the headline of the second; thanks to the community member who caught it.

How it works

  • Backbone: FINAL-Bench/Darwin-27B-RSI, full-weight trained as a decision engine.
  • Readout: the options are listed in the prompt with short codes. A readout head maps the final hidden state of one forward pass to one score per option. A softmax over the scores gives the distribution, and a fitted scalar temperature calibrates it.
  • Training data: public training and development splits of public decision benchmarks only. No test split of any benchmark was used for training.

Files

File What it is
model-*.safetensors, model.safetensors.index.json, config.json backbone weights (BF16)
readout.safetensors decision readout head
decision_config.json option codes, calibration temperature and provenance
tokenizer.json, tokenizer_config.json, chat_template.jinja tokenizer and prompt template

Usage

This repo ships the inference code in autojev/: the model.py and types.py of the autojev decision-model code (MIT, copyright notice in autojev/LICENSE), with support for a text-only backbone added. It needs torch, transformers (Qwen3.5 support), safetensors and pillow, and a GPU with space for about 54 GB of BF16 weights.

import sys
from huggingface_hub import snapshot_download

path = snapshot_download("FINAL-Bench/Darwin-27B-ZTC")
sys.path.insert(0, path)
from autojev.model import DecisionModel

model = DecisionModel(checkpoint=path)
row = {
    "state": {"ticket": "Customer was charged twice for the same order."},
    "question": {
        "type": "choice",
        "instructions": "What should support do?",
        "criteria": {"refund": "Refund the duplicate charge.", "escalate": "Send to billing.", "close": "No action."},
    },
}
probabilities = model.predict([row])[0]  # one probability per option, in criteria order

Citation

@misc{darwin27bztc2026,
  title  = {Darwin-27B-ZTC: a zero-token decision engine},
  author = {VIDRAFT and FINAL-Bench},
  year   = {2026},
  url    = {https://huggingface.co/FINAL-Bench/Darwin-27B-ZTC}
}
Downloads last month
2
Safetensors
Model size
26B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collections including FINAL-Bench/Darwin-27B-ZTC

Article mentioning FINAL-Bench/Darwin-27B-ZTC

Evaluation results

  • LocalLLaMA/typed-decisions leaderboard
  • Accuracy View evaluation results
    source
    General, zero-shot: never trained on the Typed Decisions train split, its workflows or its question schemas. Test split, 400 cases, 2,000 decisions, one request per case with the state and all five questions (README request shape), all answered, zero errors. Accuracy = agreement with the gold label; KL = KL(gold || prediction); Brier summed over options (scorer reproduces the README Uniform reference). One forward pass per question, no generated tokens.
    0.74 *
  • Kl From Gold View evaluation results
    source
    General, zero-shot: never trained on the Typed Decisions train split, its workflows or its question schemas. Test split, 400 cases, 2,000 decisions, one request per case with the state and all five questions (README request shape), all answered, zero errors. Accuracy = agreement with the gold label; KL = KL(gold || prediction); Brier summed over options (scorer reproduces the README Uniform reference). One forward pass per question, no generated tokens.
    0.2 *
  • Brier View evaluation results
    source
    General, zero-shot: never trained on the Typed Decisions train split, its workflows or its question schemas. Test split, 400 cases, 2,000 decisions, one request per case with the state and all five questions (README request shape), all answered, zero errors. Accuracy = agreement with the gold label; KL = KL(gold || prediction); Brier summed over options (scorer reproduces the README Uniform reference). One forward pass per question, no generated tokens.
    0.1 *