Instructions to use chaoliangUNSW/Jev-Style-Cascade-9B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- jev-style
How to use chaoliangUNSW/Jev-Style-Cascade-9B with jev-style:
pip install "jev-style[torch]"
from jev_style import JevStyle, noul, choice js = JevStyle.from_pretrained("chaoliangUNSW/Jev-Style-Cascade-9B") out = js.decide("I was charged twice for one order.", { "billing": noul("This message is about billing."), "team": choice("Which team should handle it?", ["billing", "shipping", "tech"]), }) print(out["answers"]["team"]["choice"]) - Notebooks
- Google Colab
- Kaggle
Jev-Style-Cascade-9B
Website: jevstyle.com
One /v1/systemone model built from two: Jev-Style-2B-Decision-v3 answers every question first, and only the questions it is not confident about go to JevK5-9B (by alibiserikbay, Apache-2.0). This repository holds no new weights: it is the frozen cascade (cascade.json), how its threshold was chosen, and its evaluation. Both models are downloaded from their own repositories at pinned revisions.
| Jev-Style-Cascade-9B | |
|---|---|
| JevBench public items (231) | 87.9 % (203 / 231) — Jev 1.13.0's published figure on the same items is 86.6 % |
| Questions the 2B settles on its own | 53 % of JevBench items, 43 % of our 4,992-question calibration set |
| Calibration error on JevBench (ECE) | 0.026 |
| Decision Index 0.2.1 (offline forecast) | 46.08, from 34.83 for the 2B alone |
| Latency, RTX 5090 (median per request) | 29 ms on JevBench items, 38 ms on the calibration set |
| Context | up to 25,600 tokens per request (tested to 25,275) |
JevBench: harness github.com/fstandhartinger/jevbench @ bb05a335 (v1.4.2.2 + 5 commits), the 231 public items, typesafe adapter against our server, one run on 2026-10-02 (UTC). This is our own run on the public items only, not a board row; the board's ranking also uses sealed items. The Jev 1.13.0 figure is public_accuracy 0.8658 in the harness's results/v1.4.2/jevbench-v1.4.2-results.json. Decision Index: see Decision Index.
Quick start
It needs one NVIDIA GPU with 32 GB (both models stay resident; see Memory), PyTorch, jev-style 0.4.0 or later, and JevK5's runtime allebee/jevk5, which is not on PyPI:
pip install "jev-style[torch]>=0.4.0" "jevk5[fast] @ git+https://github.com/allebee/jevk5@v0.3.3"
export PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True
jev-style download --cascade cascade-9b # both tiers; JevK5-9B is checked against its SHA256SUMS
jev-style serve --cascade cascade-9b --backend torch --device cuda
curl -s localhost:8765/v1/systemone -H 'content-type: application/json' -d '{
"state": "I was charged twice for my order, please refund one.",
"questions": {
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "payments and refunds", "tech": "bugs and outages"}},
"refund": {"type": "noul", "instructions": "The customer asks for money back."}}}'
From Python: from jev_style import JevStyle; JevStyle(cascade="cascade-9b").decide(state, questions). Every answer has the usual systemone shape plus tier (1 = the 2B, 2 = JevK5-9B) and tier_model; timing.tiers reports what each tier did.
How it works
- The 2B (PyTorch float32, CUDA graphs) reads the state once and answers every question in one call.
- A question's confidence is
(k · p_max − 1) / (k − 1)for k options (|2 · P(true) − 1|for true/false). Questions at τ = 0.75 or above keep the 2B's answer. - The rest go to JevK5-9B in one call with the same state; its answer is final. If the 2B fails on a question (for example an input it cannot take), JevK5-9B answers it; if JevK5-9B fails, the 2B's answer is used.
JevK5-9B runs in the same process through its own runtime (the code jevk5-serve uses, at v0.3.3), with its calibration temperatures from jevk5_config.json; the cascade refuses to load if that file is missing, since JevK5 would otherwise run uncalibrated.
How Ï„ was chosen
The rule was written down before any result existed (validation/PREDECLARATION.en.md; the binding Chinese original and its sha256 are next to it):
- Data: a 4,992-question calibration set from 29 public datasets (human or gold labels, about a quarter non-English, states up to 12K tokens). It was de-duplicated against Decision Index and JevBench, and contains nothing that either model was trained on according to their cards. The source list and licences are in validation/calibration_manifest.json.
- Constraint: pick the Ï„ with the lowest expected latency whose calibration accuracy stays within 1.0 point of JevK5-9B alone.
| Calibration set (4,992 questions) | Accuracy |
|---|---|
| Jev-Style-Cascade-9B (Ï„ = 0.75) | 73.1 % |
| JevK5-9B alone | 74.1 % |
| Jev-Style-2B-Decision-v3 alone | 66.2 % |
The 2B settled 42.8 % of the calibration questions; JevK5-9B was called for 55.5 % of the requests. The full search output is validation/freeze_result.json. An earlier run without the 2B's CUDA-graph runtime failed the pre-declared speed gate and is kept as validation/freeze_result_first_run_without_cuda_graphs.json.
Results
JevBench public items
| All (231) | Easy (48) | Standard (72) | Hard (111) | ECE | p50 latency | |
|---|---|---|---|---|---|---|
| Jev-Style-Cascade-9B | 203 | 48 | 70 | 85 | 0.026 | 29.0 ms |
| JevK5-9B alone (our run) | 203 | 48 | 69 | 86 | 0.034 | 28.1 ms |
| Jev-Style-2B-Decision-v3 alone (CUDA graphs) | 170 | 48 | 69 | 53 | 0.062 | 15.5 ms |
The 2B answered 123 of the 231 items itself. Per-item results are in validation/jevbench/.
Decision Index (offline forecast)
46.08 on Decision Index 0.2.1, against 34.83 for the 2B alone. This is a forecast, not a board entry: the cascade's answers were composed at the frozen τ from two public per-row result sets — our 2B's run (chaoliangUNSW/decision-index-results) and JevK5-9B v0.3.3's run (alibiserikbay/decision-index-results) — and scored with the official kit (apolinario/decision-index @ 87d4650) on the full rebuilt suite. Composing the same way reproduced the live cascade exactly on all 231 JevBench items. Scores: validation/decision_index_0.2.1_offline_forecast.scores.json.
Latency
| RTX 5090, in-process, one request at a time | median | mean | p95 |
|---|---|---|---|
| Calibration set (1,000 requests, up to 12K tokens) | 37.5 ms | 139.0 ms | 673.6 ms |
| JevBench public items (231, one question each) | 29.0 ms | 477.4 ms |
torch 2.14.1+cu130, transformers 5.18.0, flash-linear-attention 0.5.2, jevk5 0.3.3; the models are loaded and their CUDA graphs recorded before timing (start-up 47 s with the weights cached, including the SHA256SUMS check of JevK5-9B). A 25,275-token request with two questions took 2.4 s.
Memory
On the RTX 5090 (32 GB) the cascade used 29.7 GB after start-up and at most 31.8 GB while serving the calibration set, a 25,275-token request included, with no out-of-memory errors. Keep the two models in one process (as jev-style serve --cascade cascade-9b does) and set PYTORCH_CUDA_ALLOC_CONF=expandable_segments:True; a smaller GPU will not hold both models.
Scope and notes
- JevK5-9B is English-focused; about a quarter of the calibration set was non-English.
- JevK5-9B's model card says part of its training labels were generated by OpenAI's GPT-6 Luna under OpenAI's terms; check that those terms suit your use.
- No output of TypeSafe's Jev was used to train either model, to choose Ï„, or to evaluate the cascade; the Jev figure above is the one JevBench publishes.
Credits and licences
- Jev-Style-2B-Decision-v3 (Apache-2.0), fine-tuned from Qwen/Qwen3.5-2B (Apache-2.0).
- JevK5-9B by alibiserikbay (Apache-2.0) and its runtime allebee/jevk5 (Apache-2.0).
Neither model's weights are redistributed here. This cascade is not affiliated with, endorsed by or connected to TypeSafe or Jev, the JevK5 author, or Alibaba Cloud / the Qwen team. "Jev-style" only describes the kind of model.
Citation
@misc{jevstylecascade9b2026,
title = {Jev-Style-Cascade-9B: a confidence cascade of Jev-Style-2B-Decision-v3 and JevK5-9B},
author = {chaoliangUNSW},
year = {2026},
url = {https://huggingface.co/chaoliangUNSW/Jev-Style-Cascade-9B}
}
- Downloads last month
- -