Jeeves-9B

A Jev-like decision model that reasons before it decides. Give it a state (text or JSON) and yes/no, multiple-choice or rating questions; it writes a reasoning chain per question and returns a calibrated probability for every option.

This repository holds the fused weights (Qwen3.5-9B with the LoRA merged), the pointer head, the fitted temperature, and two speculative-decoding drafters. It loads with the Jeeves code, not with transformers: the pointer head and the prompt format are part of the model.

Files

file contents
model-*.safetensors, config.json fused Qwen3.5-9B weights, bf16
head.pt pointer head (query and key projections, 256-dim)
export.json temperature (1.859), prompt format, training step and config
tokenizer.json, vocab.json, merges.txt, tokenizer_config.json, chat_template.jinja Qwen3.5 tokenizer
drafter_k4.safetensors block-4 drafter, the serving default, bf16
drafter_k8.safetensors block-8 drafter, faster for single-question requests, bf16
drafter_k4.config.json, drafter_k8.config.json drafter training settings
LICENSE Apache-2.0, inherited from Qwen3.5-9B

Usage

hf download PostHog/jeeves --local-dir jeeves-weights
git clone https://github.com/PostHog/jeeves && cd jeeves
python -m inference.serve --model ../jeeves-weights --drafter ../jeeves-weights/drafter_k4.safetensors --port 8009
curl -s localhost:8009/v1/systemone -H 'content-type: application/json' -d '{
  "state": "I was charged twice for one order.",
  "questions": {
    "billing": {"type": "noul", "instructions": "Is this about billing?"},
    "tone": {"type": "choice", "instructions": "What is the customer'"'"'s tone?", "criteria": {"calm": null, "frustrated": null, "angry": null}}
  }}'

The request and response follow Jev's /v1/systemone format. An optional options object sets think, max_think (truncate chains), nothink_threshold (skip thinking when the no-think answer is already confident) and return_reasoning. Serving needs a CUDA GPU; the FP8 kernel needs Hopper, and --no-fp8 serves in bf16.

Results

Accuracy with thinking, greedy, 2,560-token cap. Kev-9B and Jev numbers are the ones Kev publishes. JevBench uses the same 231 public items for every model (the sealed judge tier is not included); the other rows use the same sources with different items.

Kev-9B Jev Jeeves-9B
Test overall (out-of-domain and held-out) 0.822 0.857 0.889
MMLU-Pro and buried state 0.579 0.800 0.746
JevBench, public tiers 0.715* 0.866 0.935
JevBench hard 0.451* 0.730 0.865

* Kev-8B (Qwen3); no Kev-9B JevBench result is published.

Without thinking the model scores 0.804 on the Jeeves test split, against 0.840 with it. With the fitted temperature the no-think path has a top-label calibration error of 0.021 and 1.3% confident errors (wrong at p ≥ 0.9).

Serving on one H100, 325 dev questions:

setting accuracy median / p90 latency
full thinking 0.825 3.3 s / 17.1 s
max_think 768, nothink_threshold 0.9 0.806 2.0 s / 5.6 s
no thinking 0.775 about 0.3 s

Training

  • SFT. LoRA r=16 on every projection plus the pointer head, 2 epochs on 19,126 questions from 12 public datasets and synthetic policy data; half the questions carry a reasoning chain sampled from the base model.
  • CISPO. 9,992 RL questions, 8 rollouts each, up to 2,560 thinking tokens. Reward is the probability of the correct option, discounted by up to 10% for long chains. The released checkpoint is step 402 of a 624-step schedule.
  • Calibration. One softmax temperature fitted on the dev set.
  • Drafters. Diffusion views of the frozen model, inspired by Orthrus, distilled by KL divergence on chains sampled from this model. At one question the block-4 drafter decodes 1.6× faster than graphed greedy decoding and the block-8 drafter 1.76×.

A from-scratch reproduction with the released code matched this checkpoint within noise. The training code, dataset build scripts and pinned dataset revisions are in the GitHub repository.

Limitations

  • Knowledge questions trail Jev (MMLU 0.793 vs 0.900, MMLU-Pro 0.739 vs 0.840).
  • Thinking is slow at the tail (17 s at p90 with full chains); use max_think and nothink_threshold when latency matters.
  • About a third of full-length chains hit the token cap without closing; accuracy on those items is lower.
  • Evaluated on English inputs only.

License

The weights are derived from Qwen3.5-9B and are released under its Apache-2.0 license (LICENSE). The Jeeves code is MIT licensed.

Citation

@software{waltz2026jeeves,
  author = {Waltz, Nicholas P.},
  title  = {Jeeves: Reasoning Improves Jev-like Decisions},
  year   = {2026},
  url    = {https://github.com/PostHog/jeeves},
  note   = {Qwen3.5-9B decision model trained with SFT and CISPO, with a block-4 diffusion drafter}
}
Downloads last month
61
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for PostHog/jeeves

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(967)
this model
Quantizations
1 model