Jeeves-9B
A Jev-like decision model that reasons before it decides. Give it a state (text or JSON) and yes/no, multiple-choice or rating questions; it writes a reasoning chain per question and returns a calibrated probability for every option.
This repository holds the fused weights (Qwen3.5-9B with the LoRA merged), the pointer head, the fitted temperature, and two speculative-decoding drafters. It loads with the Jeeves code, not with transformers: the pointer head and the prompt format are part of the model.
Files
| file | contents |
|---|---|
model-*.safetensors, config.json |
fused Qwen3.5-9B weights, bf16 |
head.pt |
pointer head (query and key projections, 256-dim) |
export.json |
temperature (1.859), prompt format, training step and config |
tokenizer.json, vocab.json, merges.txt, tokenizer_config.json, chat_template.jinja |
Qwen3.5 tokenizer |
drafter_k4.safetensors |
block-4 drafter, the serving default, bf16 |
drafter_k8.safetensors |
block-8 drafter, faster for single-question requests, bf16 |
drafter_k4.config.json, drafter_k8.config.json |
drafter training settings |
LICENSE |
Apache-2.0, inherited from Qwen3.5-9B |
Usage
hf download PostHog/jeeves --local-dir jeeves-weights
git clone https://github.com/PostHog/jeeves && cd jeeves
python -m inference.serve --model ../jeeves-weights --drafter ../jeeves-weights/drafter_k4.safetensors --port 8009
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
"state": "I was charged twice for one order.",
"questions": {
"billing": {"type": "noul", "instructions": "Is this about billing?"},
"tone": {"type": "choice", "instructions": "What is the customer'"'"'s tone?", "criteria": {"calm": null, "frustrated": null, "angry": null}}
}}'
The request and response follow Jev's /v1/systemone format. An optional options object sets think, max_think (truncate chains), nothink_threshold (skip thinking when the no-think answer is already confident) and return_reasoning. Serving needs a CUDA GPU; the FP8 kernel needs Hopper, and --no-fp8 serves in bf16.
Results
Accuracy with thinking, greedy, 2,560-token cap. Kev-9B and Jev numbers are the ones Kev publishes. JevBench uses the same 231 public items for every model (the sealed judge tier is not included); the other rows use the same sources with different items.
| Kev-9B | Jev | Jeeves-9B | |
|---|---|---|---|
| Test overall (out-of-domain and held-out) | 0.822 | 0.857 | 0.889 |
| MMLU-Pro and buried state | 0.579 | 0.800 | 0.746 |
| JevBench, public tiers | 0.715* | 0.866 | 0.935 |
| JevBench hard | 0.451* | 0.730 | 0.865 |
* Kev-8B (Qwen3); no Kev-9B JevBench result is published.
Without thinking the model scores 0.804 on the Jeeves test split, against 0.840 with it. With the fitted temperature the no-think path has a top-label calibration error of 0.021 and 1.3% confident errors (wrong at p ≥ 0.9).
Serving on one H100, 325 dev questions:
| setting | accuracy | median / p90 latency |
|---|---|---|
| full thinking | 0.825 | 3.3 s / 17.1 s |
max_think 768, nothink_threshold 0.9 |
0.806 | 2.0 s / 5.6 s |
| no thinking | 0.775 | about 0.3 s |
Training
- SFT. LoRA r=16 on every projection plus the pointer head, 2 epochs on 19,126 questions from 12 public datasets and synthetic policy data; half the questions carry a reasoning chain sampled from the base model.
- CISPO. 9,992 RL questions, 8 rollouts each, up to 2,560 thinking tokens. Reward is the probability of the correct option, discounted by up to 10% for long chains. The released checkpoint is step 402 of a 624-step schedule.
- Calibration. One softmax temperature fitted on the dev set.
- Drafters. Diffusion views of the frozen model, inspired by Orthrus, distilled by KL divergence on chains sampled from this model. At one question the block-4 drafter decodes 1.6× faster than graphed greedy decoding and the block-8 drafter 1.76×.
A from-scratch reproduction with the released code matched this checkpoint within noise. The training code, dataset build scripts and pinned dataset revisions are in the GitHub repository.
Limitations
- Knowledge questions trail Jev (MMLU 0.793 vs 0.900, MMLU-Pro 0.739 vs 0.840).
- Thinking is slow at the tail (17 s at p90 with full chains); use
max_thinkandnothink_thresholdwhen latency matters. - About a third of full-length chains hit the token cap without closing; accuracy on those items is lower.
- Evaluated on English inputs only.
License
The weights are derived from Qwen3.5-9B and are released under its Apache-2.0 license (LICENSE). The Jeeves code is MIT licensed.
Citation
@software{waltz2026jeeves,
author = {Waltz, Nicholas P.},
title = {Jeeves: Reasoning Improves Jev-like Decisions},
year = {2026},
url = {https://github.com/PostHog/jeeves},
note = {Qwen3.5-9B decision model trained with SFT and CISPO, with a block-4 diffusion drafter}
}
- Downloads last month
- 61