Kodiak-v0.4-1B, accuracy mode
The most accurate and best-calibrated Kodiak: three Kodiak-v0.4-1B models answer every question and their calibrated answers are averaged. It costs about 3× the compute of the single model. Kodiak is built by Cortex Agent LLC. You give it a state (text, a list of texts, or JSON) and typed questions. It returns calibrated choice, score or "can't tell" answers in one forward pass. Use it to automate routine "read this and decide" work: routing, triage, guardrails and checks. Send the cases it isn't sure about to a person or an LLM.
Built on Ettin-encoder-1B (Johns Hopkins, MIT license), with 1.04 billion parameters. Code, docs and the full public build log: https://github.com/grizzlypeaksoftware/kodiak
# pip install "kodiak-s1[infer] @ git+https://github.com/grizzlypeaksoftware/kodiak"
from kodiak_s1.hub import Kodiak
kodiak = Kodiak.from_pretrained("cortex-agent-llc/kodiak-v0.4-1b-accuracy") # loads all three members
# Guardrail against your own written policy
kodiak.decide(
{"policy": "1. No selling outside the marketplace. 2. All prices must be listed in USD. 3. Never post a personal phone number.",
"message": "Road bike for sale, $250. Call me on 555-201-7788 if you want to see it."},
[{"type": "choice", "id": "rule", "text": "Which rule does the message break?",
"labels": ["rule 1", "rule 2", "rule 3", "no rule"]}],
)
# -> rule: "rule 3" (0.77)
# Refund eligibility under a written policy
kodiak.decide(
{"policy": "1. Refunds within 30 days of delivery. 2. Item must be unused and in its original packaging. 3. Sale items are final, no refunds.",
"request": "I bought these headphones at full price 45 days ago and they arrived unopened. Can I get a refund?"},
[{"type": "choice", "id": "refund", "text": "Is the customer eligible for a refund under this policy?",
"labels": ["yes, eligible", "no, not eligible", "can't decide without more information"]}],
)
# -> refund: "no, not eligible" (1.00)
For the fastest answers, use the single model: Kodiak-v0.4-1B. The example outputs above are from the single model; accuracy mode's probabilities differ slightly. In accuracy mode, the abstain threshold (0.6) was tuned on validation data only. On the contrastive tests accuracy mode scores: refund 0.80 (first test) and 0.64 (cue-hard); step safety 0.40 and 0.33.
What's new in v0.4
After v0.3, a Hugging Face user showed that two skills were answering from keywords. We built contrastive tests (the same case twice, one detail changed, a different correct answer; a keyword shortcut can't get both right), then trained on contrast groups: one shared field written under every answer, with the versions forced to use the same linking words ("but", "if", "before"), so only the facts decide. Scores below are the share of pairs where both versions are answered correctly (three-run means):
| Contrastive test | Kodiak-v0.3-1B | Kodiak-v0.4-1B |
|---|---|---|
| Refund eligibility, first test (50 pairs) | 0.22 | 0.79 |
| Refund eligibility, cue-hard pairs (128 pairs a changed-words reader gets wrong) | 0.27 | 0.58 |
| Agent step safety, first test (50 pairs) | 0.05 | 0.35 |
| Agent step safety, cue-hard pairs (54) | 0.13 | 0.29 |
Refund is a real fix. Step safety improved but is not fixed (see Known limits). The case the user reported, git push origin main
while three tests fail, now gets "ask the user first" (0.95); v0.3 said "yes, safe" (0.97).
Measured on real data the model never trained on (chance-corrected skill, 0 = random guessing; three-run means):
| Real-data test | Kodiak-v0.3-1B | Kodiak-v0.4-1B | v0.4 accuracy mode |
|---|---|---|---|
| RAGBench: is a long answer supported by its sources? (CC BY 4.0) | +0.41 | +0.39 | +0.40 |
| MT-Bench: which answer did expert humans prefer? (CC BY 4.0) | +0.29 | +0.30 | +0.31 |
| SemEval-2016: stance of a tweet toward a target (MIT) | +0.37 | +0.38 | +0.37 |
| Average | 0.36 | 0.36 | 0.36 |
How it compares (frozen eval set v0.2, choice questions)
| Kodiak-v0.3-1B | Kodiak-v0.4-1B | v0.4 accuracy mode | Qwen3-8B (LLM) | |
|---|---|---|---|---|
| Never-seen tasks, forced accuracy | 0.687 ± 0.010 | 0.685 ± 0.011 | 0.696 | 0.688 |
| Familiar tasks | 0.879 | 0.880 | 0.888 | 0.710 |
| Ranks its own mistakes last (never-seen) | 54.5% | 54.5% | 56.1% | 14.7% |
| Calibration error (never-seen; lower is better)¹ | 0.076 | 0.070 | 0.059 | 0.293 |
| When it says "can't tell", it's right | 0.93 | 0.94 | 0.90 | – |
| Latency (GPU, one request) | ~38 ms | ~38 ms | ~3× | ~1,500 ms |
Kodiak-v0.4-1B figures with ± are the mean of three training runs. This checkpoint is the seed-1 run, chosen by validation loss and never by the eval set. Its own scores are 0.676 never-seen forced, 0.873 familiar, 0.061 never-seen calibration error and 0.972 "can't tell" precision. "Never-seen" means tasks and label sets the model was never trained on (zero-shot). v0.4 does not improve never-seen accuracy; its gains are refund eligibility, calibration and shortcut-proof testing.
Ranking is what a confidence threshold relies on: sort answers by confidence and see how much of the gap between a random order and the perfect order (all mistakes last) the model closes. ¹ The calibration comparison is raw. Qwen3-8B is mostly overconfident by a constant amount, so a fitted recalibration brings it from 0.293 to 0.178.
Known limits (read before using)
- Agent step safety is not reliable. Do not use it as a safety control. It now reads the task (deleting the task changes 28% of answers, up from 4% in v0.3), but on pairs built so that hedge words don't give the answer away it gets both versions right only 29% of the time, and an unrelated sentence added to the input changes 21% of its answers.
- Refund eligibility is much better, not perfect. In our own spot checks it still misses some single conditions (a "sale items are final" rule) with low confidence, and rarely says "can't decide" when the request leaves out a fact the policy needs. Use a confidence threshold.
- Sarcasm has no signal on real tweets yet. Our synthetic sarcasm test can be passed by a word counter. On the Decision Index (v0.2, same training recipe for sarcasm), picking the sarcastic tweet of a pair scored 0.505, a coin flip.
- Pairwise judging has some position bias. Swapping the two answers changes 18% of its verdicts on MT-Bench (v0.3: 26%).
- Option wording still matters. Reworded options change the answer about a third of the time on never-seen tasks. Keep options short and distinct, and test a few wordings on your own data. On a 100-item wording-trap test it is right 89% of the time.
- New kinds of task are its weakest area. It is good at the kinds of decision it was trained on and much weaker on genuinely new ones.
- Implied violations are harder than stated ones; treat low-confidence guardrail answers as "send to a person".
- Arithmetic (e.g. does an invoice total match its lines?) is near chance and wasn't trained.
- Long inputs. The position limit is about 8,000 tokens (the state, every question and every option together). Longer requests are refused rather than truncated.
- Ratings ("how urgent is this?") are rough. Prefer choice questions.
- Speed. About 38 ms per request on a GPU. On a CPU, expect a few hundred milliseconds.
- Validate it on your own data. Don't use it for decisions about people without human review. Teach it your own decisions with the fine-tuning kit in the repository (a CSV in, a before/after report out).
License and citation
Apache-2.0 (weights and code). Base model: Ettin-encoder-1B (MIT). Training-data licenses are listed in data/LICENSES.md in the repository.
All synthetic training data was written and checked by open-weight models; no closed-model outputs and no evaluation data are in training.
Cortex Agent LLC / Grizzly Peak Software, 2026.
Model tree for cortex-agent-llc/kodiak-v0.4-1b-accuracy
Base model
jhu-clsp/ettin-encoder-1b