JPT-4B

License JevBench v1.4 public Decision Index 0.2.1 llm2jev

JPT-0.8B ยท JPT-4B ยท JPT-9B ยท JPT-35B-A3B ยท llm2jev

2048 Snake Tetris
2048 Snake Tetris

Each move is one choice question answered in one forward pass (30 moves each, seed 0, SGLang, recorded headless).

What is JPT-4B

JPT-4B is a fast, open decision model: give it a situation and typed questions, get a calibrated probability for every option from one forward pass. No generated explanation, no reasoning tokens โ€” latency is one prefill.

It implements the typed-decision interface introduced by Jev from TypeSafe AI [1]: a caller sends a state plus questions, and each question is one of three types. JPT is an independent model, not derived from Jev and not trained on Jev outputs; it is an open alternative behind the same interface.

Question type What it answers Options
choice pick one 2โ€“255 labels
score a level on an ordered scale the scale's levels
noul yes / no true, false

Built on Qwen/Qwen3.5-4B: a LoRA fine-tune merged into full weights. The vision tower is unchanged, so screenshots, photos and video frames work next to the text state.

Probabilities use one temperature T = 1.036, fit once on a held-out split โ€” never per benchmark. Read them as well-ordered confidence; ECE varies by task (0.02โ€“0.19).

โฏโฏ Benchmarks

โฏ JevBench v1.4.0

JevBench [2] scores general typed decisions; JPT-4B has the highest public accuracy on it of any system in the v1.4 results. v1.4.0 has 231 public items (frozen since v1.2) and 308 sealed items that only the maintainer can run, so JPT-4B has no official v1.4 score yet. Our runner gives JevK5 0.861 on the public items against its official 0.853.

JevBench v1.4.0 public accuracy

Table
System Params Public accuracy (231) Sealed accuracy (308) v1.4 score
JPT-35B-A3B (ours) 35B-A3B 0.892 pending pending
JPT-4B 4B 0.879 pending pending
Jev 1.13.0 (TypeSafe AI, API) closed 0.866 0.367 63.3
Winnow-12B Q8 12B 0.857 0.331 55.6
JevK5 v0.2.0 27B 0.853 0.331 62.0
decider-35b-a3b 35B-A3B 0.831 0.315 41.2
Hopper โ€” 0.823 0.341 59.4
openjev 4B v5ยน 4B 0.814 โ€” โ€”
SemIf, formerly OpenJev (Qwen3.5-4B) 4B 0.810 0.263 47.7
local-jev Qwen3.5-4B 4B 0.805 0.260 46.8
reflex 4B 4B 0.792 0.282 54.0
Qwen3.5-4B, same prompt, zero-shot (our run) 4B 0.740 โ€” โ€”
kev 4B (research preview) 4B 0.662 0.224 36.1

Source: other rows from results/v1.4/jevbench-v1.4-results.json at jevbench commit 2fa63fa (2026-09-23). ยน Not in the v1.4 results; its own card's number on the 231 public items with the benchmark's harness.

โฏ Decision Index 0.2.1

Decision Index [3] 0.2.1 (2026-09-27) is the broadest test: the full frozen suite, 38 scored benchmarks in five areas, chance-corrected โ€” JPT-4B scores 43.04, ahead of every other 4B entrant. We ran it through llm2jev over SGLang (apolinario/decision-index#7). JPT-4B is not on the live board yet, so its number is our rescore of that run with the benchmark's own kit (decision-index score --edition 0.2.1; the same command reproduces JPT-9B's board entry exactly). It was 39.54 under edition 0.2. The other rows are from the live 0.2.1 board.

Decision Index 0.2.1

Table
Model Params Decision Index 0.2.1
Jev 1.13.0 (TypeSafe AI, API) closed 57.91
JPT-35B-A3B (ours, not on the board yet) 35B-A3B 52.89
JPT-4B 4B 43.04
Hopper (G) 1.2 4B 40.77
Decider 4B 4B 40.70
JevK5 4B 38.81
NeoHorse-Jev-4B 4B 36.75
Kev 4B 4B 34.64

Source: live board data/index-v0.2.1.json (generated 2026-09-27 16:59 UTC); our run and scores.json in kirp/decision-index-results-jpt-4b (gated: it carries the suite's GPQA/HLE item text).

By area, against Jev 1.13.0 on the same items and scorer: JPT-4B is behind Jev in all five areas and ahead of it on 4 of the 38 index benchmarks.

Decision Index 0.2.1 by area

Per-area and per-benchmark skill vs Jev 1.13.0
Area Jev 1.13.0 JPT-4B
Knowledge 51.4 28.7
Language 62.0 52.5
Retrieval 55.4 45.0
Tools 75.1 57.2
Arts 37.7 25.8
Area Benchmark Jev 1.13.0 JPT-4B
Arts BPoMP 81.8 61.1
Arts ForecastBench 30.6 24.5
Arts Habermas Machine 21.5 11.3
Arts Humicroedit 23.7 19.5
Arts New Yorker 62.6 46.7
Arts POP909-CL 15.9 0.7
Arts cfcolor 28.8 17.3
Games ChessBench 9.8 6.2
Knowledge BBH 89.7 46.5
Knowledge CLadder 45.3 29.9
Knowledge CRUXEval 57.1 22.1
Knowledge GPQA Diamond 71.4 23.8
Knowledge GSM8K 75.6 66.6
Knowledge HLE 4.7 0.0
Knowledge MMLU-Pro 80.5 42.5
Knowledge MuSR 46.1 29.4
Knowledge SATA-Bench 25.4 20.7
Language ACOS 27.3 17.5
Language ANLI 62.2 43.1
Language ContractNLI 59.1 70.6
Language FinEntity 80.8 88.1
Language HellaSwag 92.7 74.7
Language NLI4CT 69.0 58.1
Language RAGTruth 51.3 46.5
Language VAST 46.9 45.1
Language WinoGrande 83.9 42.5
Language iSarcasmEval 36.3 37.9
Retrieval Amazon ESCI 43.8 32.8
Retrieval BANKING77 79.5 71.8
Retrieval BRIGHT 40.6 34.6
Retrieval CLINC150+OOS 89.2 64.3
Retrieval HoVer 45.7 21.7
Retrieval PhishNChips phishing decisions 25.1 37.8
Tools API-Bank 88.0 68.1
Tools BFCL 94.3 88.6
Tools Home appliance simulator 52.3 12.5
Tools ToolRet 59.9 56.2
Tools When2Call 74.6 52.1

Chance-corrected skill ร— 100 (0 = random, 100 = perfect). Jev's numbers are its official entry on the live board (jev-1.13.0); ours are from the same kit and suite.

โฏ Against its base model

Same prompt, same serving path, Qwen3.5-4B zero-shot vs JPT-4B: fine-tuning lifts every text benchmark and leaves images roughly where the base model was.

JPT-4B vs Qwen3.5-4B

Table, with calibration
Benchmark (version, n) What it tests JPT-4B Qwen3.5-4B zero-shot Jev 1.13.0
JevBench v1.4.0 public hard tier [2] (111) hardest general decisions 0.784 (ECE 0.068, Brier 0.318) 0.595 โ€”
Typed decisions test (ours, 2,000) in-distribution typed decisions 0.796 (ECE 0.160, Brier 0.332) 0.596 (ECE 0.171, Brier 0.559) โ€”
ANLI r1 / r3 [4] (dev) adversarial NLI 0.697 / 0.613 0.660 / 0.513 โ€”
Banking77 [5] / MASSIVE 1.1 [6] (en / de / zh) intent classification 0.757 / 0.857 / 0.833 / 0.837 0.663 / 0.733 / 0.670 / 0.703 โ€”
ScreenSpot-v2 [7] set-of-marks (300, images) GUI element grounding 0.923 (ECE 0.029) 0.903 (ECE 0.058) โ€”
Screen2Words [8] match (300, images) screenshot summarization 0.927 (ECE 0.037) 0.887 (ECE 0.030) โ€”
ERQA [9] (400, images) embodied visual reasoning 0.455 (ECE 0.150) 0.463 (ECE 0.100) โ€”
EnvBench v0.1 (ours) public / held-out [10] (skill 0โ€“100) sequential decisions in game envs 47.7 / 47.0 โ€” โ€”

Jev 1.13.0 has no official score on these splits (our own test/dev cuts, EnvBench, and the image sets), so its column is "โ€”"; its official scores on the Decision Index versions of ANLI and BANKING77 are in the per-benchmark table above. Running the Jev API on these splits would fill them.

Banking77, MASSIVE and typed rows are in-distribution (train splits in the mix, test items not). Image rows are zero-shot: no image data was trained on.

โฏโฏ Quick Start

Two pieces: an engine that holds the weights, and llm2jev (>= 0.6.1) in front of it, reading option probabilities off the engine.

โšก SGLang (recommended)

python -m sglang.launch_server --model-path kirp/jpt-4b --port 30000 \
  --context-length 32768 --mamba-scheduler-strategy extra_buffer &  # Qwen3.5's DeltaNet layers need this flag
llm2jev --model kirp/jpt-4b --backend sglang --url http://127.0.0.1:30000 --port 8080 --temperature 1.036

Tested with SGLang 0.5.9; install cuDNN 9.15+ over its pinned 9.10: pip install "sglang==0.5.9" && pip install "nvidia-cudnn-cu12>=9.15".

๐Ÿ” vLLM

vllm serve kirp/jpt-4b --max-logprobs 256 --return-tokens-as-token-ids --enable-scale-out --port 8000
llm2jev --model kirp/jpt-4b --backend vllm --url http://127.0.0.1:8000 --port 8080 --temperature 1.036

The three vLLM flags are required: without them every request is a bare HTTP 400.

๐Ÿงช No engine (quick check only)

pip install "llm2jev[hf,vision]"
llm2jev --model kirp/jpt-4b --backend hf --port 8080 --temperature 1.036

Serializes requests; for traffic use SGLang or vLLM.

๐Ÿ“จ Ask it a question

import requests
r = requests.post("http://127.0.0.1:8080/v1/systemone", json={
    "state": "Refund policy: full refund within 30 days of purchase; 50% until day 60; none after.\n"
             "Order 1182 was bought on 3 March and returned on 20 April.",
    "questions": {
        "refund": {"type": "choice", "instructions": "What refund does order 1182 get?",
                   "criteria": {"full": "Full refund", "half": "50% refund", "none": "No refund"}},
        "late":   {"type": "noul", "instructions": "Was the return made after day 30?",
                   "criteria": {"true": "Yes", "false": "No"}}}})
print(r.json()["answers"])   # each answer has the per-option probabilities

๐Ÿ–ผ๏ธ With images

state = [{"role": "user", "content": [
    {"type": "image", "image": "https://example.com/screen.png"},
    {"type": "text", "text": "Task: open the settings page. Numbered boxes mark clickable elements."}]}]
questions = {"click": {"type": "choice", "instructions": "Which box should be clicked?",
                       "criteria": {"1": None, "2": None, "3": None, "4": None, "5": None}}}

โฏโฏ Training

LoRA with a Brier loss on 49,221 typed questions, one epoch, merged into full weights.

Part What it is
Method LoRA r=16 on every attention, DeltaNet and MLP projection of the language model; vision tower untouched
Loss multi-class Brier over the option labels, on llm2jev's chat prompt with thinking disabled
Data 49,221 questions in 32,835 records; one epoch over two option-shuffled copies
Held out no item from JevBench, EnvBench held-out seeds, the Decision Index frozen suite or our typed test split; 8-gram overlap check vs JevBench
Data sources
  • public classification, NLI, QA, preference and safety datasets;
  • long legal and contract documents (ContractNLI, MAUD, LegalBench, ConditionalQA, ShARC);
  • table and numeric reasoning (TAT-QA, MultiHiertt);
  • multi-hop QA (MuSiQue, BEIR);
  • agent and tool traces (Mind2Web, AgentTraj, ToolACE);
  • community typed-decision sets;
  • oracle-labelled rollouts from 20 small game and puzzle environments;
  • programmatically generated rule-arithmetic items (dates, time zones, day counts, caps; labels computed by code);
  • 1,289 long policy / contract / regulation documents (4,471 questions) written by an LLM (GPT-6 Luna) with no JevBench item shown to it.

โฏโฏ Limitations

  • Arithmetic and dates are its weakest area: it answers from the evidence given and has no reasoning phase by design.
  • Up to 255 options are accepted; training covered up to 77 (Banking77), so beyond that quality is not established.
  • English first. Other languages come only from a few multilingual classification sets.
  • Images are zero-shot through the base vision tower.
  • Minesweeper-style belief reasoning is weak: in our demo loop it reveals cells in reading order and loses fast.

โฏโฏ References

  1. TypeSafe AI. Jev. https://typesafe.ai
  2. F. Standhartinger. JevBench, v1.4.0. https://github.com/fstandhartinger/jevbench
  3. Decision Index, v0.2.1. https://huggingface.co/spaces/multimodalart/jev-decision-index
  4. Nie et al. Adversarial NLI. ACL 2020.
  5. Casanueva et al. Efficient Intent Detection with Dual Sentence Encoders (Banking77). NLP4ConvAI 2020.
  6. FitzGerald et al. MASSIVE. ACL 2023.
  7. Wu et al. OS-Atlas (ScreenSpot-v2). 2024.
  8. Wang et al. Screen2Words. UIST 2021.
  9. Gemini Robotics Team. Gemini Robotics (ERQA). 2025.
  10. EnvBench, v0.1 (ours, frozen 2026-09-23; not yet public): programmatically solved game, planning and rule decisions with exact gold answers.

โฏโฏ License

CC BY-NC 4.0. The weights derive from Qwen3.5-4B (Apache-2.0), but some training datasets allow only non-commercial or research use, so the model is released for non-commercial use.


JPT-4B is an independent open model that implements a typed-decision interface (noul, choice and score questions answered with probabilities). It is not affiliated with, endorsed by or derived from TypeSafe AI or its Jev model, and it was not trained on Jev outputs.

Downloads last month
86
Safetensors
Model size
5B params
Tensor type
F32
ยท
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for kirp/jpt-4b

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(794)
this model

Spaces using kirp/jpt-4b 2