Instructions to use kirp/jpt-35b-a3b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use kirp/jpt-35b-a3b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("image-text-to-text", model="kirp/jpt-35b-a3b") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] pipe(text=messages)# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("kirp/jpt-35b-a3b") model = AutoModelForMultimodalLM.from_pretrained("kirp/jpt-35b-a3b", device_map="auto") messages = [ { "role": "user", "content": [ {"type": "image", "url": "https://huggingface.co/datasets/huggingface/documentation-images/resolve/main/p-blog/candy.JPG"}, {"type": "text", "text": "What animal is on the candy?"} ] }, ] inputs = processor.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=40) print(processor.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use kirp/jpt-35b-a3b with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "kirp/jpt-35b-a3b" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kirp/jpt-35b-a3b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker
docker model run hf.co/kirp/jpt-35b-a3b
- SGLang
How to use kirp/jpt-35b-a3b with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "kirp/jpt-35b-a3b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kirp/jpt-35b-a3b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "kirp/jpt-35b-a3b" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "kirp/jpt-35b-a3b", "messages": [ { "role": "user", "content": [ { "type": "text", "text": "Describe this image in one sentence." }, { "type": "image_url", "image_url": { "url": "https://cdn.britannica.com/61/93061-050-99147DCE/Statue-of-Liberty-Island-New-York-Bay.jpg" } } ] } ] }' - Docker Model Runner
How to use kirp/jpt-35b-a3b with Docker Model Runner:
docker model run hf.co/kirp/jpt-35b-a3b
JPT-35B-A3B
JPT-0.8B · JPT-4B · JPT-9B · JPT-35B-A3B · llm2jev
Each move is one choice question answered in one forward pass (30 moves each, seed 0, SGLang, recorded headless).
What is JPT-35B-A3B
JPT-35B-A3B is the largest JPT: an open decision model that takes a situation and typed questions and returns a calibrated probability for every option from one forward pass. No generated explanation, no reasoning tokens — latency is one prefill, and only 3B of its 35B parameters are active per token.
It implements the typed-decision interface introduced by Jev from TypeSafe AI [1]: a caller sends a state plus questions, and each question is one of three types. JPT is an independent model, not derived from Jev and not trained on Jev outputs; it is an open alternative behind the same interface.
| Question type | What it answers | Options |
|---|---|---|
choice |
pick one | 2–255 labels |
score |
a level on an ordered scale | the scale's levels |
noul |
yes / no | true, false |
Built on Qwen/Qwen3.5-35B-A3B, a mixture-of-experts model (3B active parameters per token): a LoRA fine-tune merged into full weights, with the same recipe and data as JPT-4B and JPT-9B. The vision tower is unchanged, so screenshots, photos and video frames work next to the text state.
Probabilities use one temperature T = 1.087, fit once on a held-out split — never per benchmark. Read them as well-ordered confidence; ECE varies by task (0.03–0.17).
❯❯ Benchmarks
❯ JevBench v1.4.0
JevBench [2] scores general typed decisions; JPT-35B-A3B gets 206 of the 231 public items, above every system in the v1.4 results. v1.4.0 has 231 public items (frozen since v1.2) and 308 sealed items that only the maintainer can run, so JPT-35B-A3B has no official v1.4 score yet.
Table
| System | Params | Public accuracy (231) | Sealed accuracy (308) | v1.4 score |
|---|---|---|---|---|
| JPT-35B-A3B | 35B-A3B | 0.892 | pending | pending |
| JPT-4B (ours) | 4B | 0.879 | pending | pending |
| Jev 1.13.0 (TypeSafe AI, API) | closed | 0.866 | 0.367 | 63.3 |
| Winnow-12B Q8 | 12B | 0.857 | 0.331 | 55.6 |
| JPT-9B (ours) | 9B | 0.853 | pending | pending |
| JevK5 v0.2.0 | 27B | 0.853 | 0.331 | 62.0 |
| Qwen3.5-35B-A3B, same prompt, zero-shot (our run) | 35B-A3B | 0.844 | — | — |
| decider-35b-a3b | 35B-A3B | 0.831 | 0.315 | 41.2 |
| Hopper | — | 0.823 | 0.341 | 59.4 |
By tier: easy 48/48, standard 71/72, hard 87/111. Latency on our runner (tensor parallel 4 on A100 40 GB, no other load on that replica): p50 56 ms, p95 103 ms.
Source: other rows from results/v1.4/jevbench-v1.4-results.json at jevbench commit 2fa63fa (2026-09-23).
❯ Decision Index 0.2.1
Decision Index [3] (0.2.1, 2026-09-27) is the broadest test: 38 scored benchmarks in five areas, chance-corrected. JPT-35B-A3B scores 52.89, 5.8 points above Decider 35B-A3B on the same base model and 6.0 above JPT-9B. We ran the full frozen suite (150,759 requests, every one answered) through llm2jev over SGLang and submitted it as apolinario/decision-index#7; it is not on the live board until the maintainer validates it.
Table
| Model | Params | Decision Index 0.2.1 |
|---|---|---|
| Jev 1.13.0 (TypeSafe AI, API) | closed | 57.91 |
| Surogate Rune 26B-A4B v3 | 26B-A4B | 57.44 |
| Decider chat · Gemma-4-31B | 31B | 57.33 |
| AutoJev-27B | 27B | 56.40 |
| simple-jev · Qwen3.8-27B | 27B | 55.74 |
| Jebadiah 27B | 27B | 54.67 |
| Eikos-27B | 27B | 53.13 |
| JPT-35B-A3B | 35B-A3B | 52.89 |
| reflex 27B | 27B | 52.16 |
| Decider chat · Qwen3.6-27B | 27B | 51.35 |
| Winnow-12B | 12B | 50.02 |
| Decider 35B-A3B | 35B-A3B | 47.11 |
| JPT-9B (ours) | 9B | 46.89 |
Source: full run and scores.json in kirp/decision-index-results-jpt-35b-a3b
(gated: it carries the suite's GPQA/HLE item text). Other rows: live board data/index-v0.2.1.json (generated 2026-09-27 16:59 UTC).
By area, against Jev 1.13.0 on the same items and scorer: JPT-35B-A3B is level with Jev on language and retrieval, behind it on knowledge, tools and arts, and ahead of it on 7 of the 38 index benchmarks.
Per-area and per-benchmark skill vs Jev 1.13.0
| Area | Jev 1.13.0 | JPT-35B-A3B |
|---|---|---|
| Knowledge | 51.4 | 37.0 |
| Language | 62.0 | 61.9 |
| Retrieval | 55.4 | 55.7 |
| Tools | 75.1 | 69.5 |
| Arts | 37.7 | 34.8 |
| Area | Benchmark | Jev 1.13.0 | JPT-35B-A3B |
|---|---|---|---|
| Arts | BPoMP | 81.8 | 80.4 |
| Arts | ForecastBench | 30.6 | 25.9 |
| Arts | Habermas Machine | 21.5 | 18.2 |
| Arts | Humicroedit | 23.7 | 24.7 |
| Arts | New Yorker | 62.6 | 60.5 |
| Arts | POP909-CL | 15.9 | 11.0 |
| Arts | cfcolor | 28.8 | 24.6 |
| Games | ChessBench | 9.8 | 8.9 |
| Knowledge | BBH | 89.7 | 60.8 |
| Knowledge | CLadder | 45.3 | 38.0 |
| Knowledge | CRUXEval | 57.1 | 48.8 |
| Knowledge | GPQA Diamond | 71.4 | 32.0 |
| Knowledge | GSM8K | 75.6 | 67.5 |
| Knowledge | HLE | 4.7 | 0.0 |
| Knowledge | MMLU-Pro | 80.5 | 56.8 |
| Knowledge | MuSR | 46.1 | 35.3 |
| Knowledge | SATA-Bench | 25.4 | 21.6 |
| Language | ACOS | 27.3 | 18.8 |
| Language | ANLI | 62.2 | 57.2 |
| Language | ContractNLI | 59.1 | 73.4 |
| Language | FinEntity | 80.8 | 90.3 |
| Language | HellaSwag | 92.7 | 89.0 |
| Language | NLI4CT | 69.0 | 66.0 |
| Language | RAGTruth | 51.3 | 49.1 |
| Language | VAST | 46.9 | 52.8 |
| Language | WinoGrande | 83.9 | 67.3 |
| Language | iSarcasmEval | 36.3 | 49.3 |
| Retrieval | Amazon ESCI | 43.8 | 41.5 |
| Retrieval | BANKING77 | 79.5 | 77.2 |
| Retrieval | BRIGHT | 40.6 | 37.2 |
| Retrieval | CLINC150+OOS | 89.2 | 85.0 |
| Retrieval | HoVer | 45.7 | 35.1 |
| Retrieval | PhishNChips phishing decisions | 25.1 | 51.5 |
| Tools | API-Bank | 88.0 | 85.4 |
| Tools | BFCL | 94.3 | 94.8 |
| Tools | Home appliance simulator | 52.3 | 35.2 |
| Tools | ToolRet | 59.9 | 57.9 |
| Tools | When2Call | 74.6 | 65.9 |
Chance-corrected skill × 100 (0 = random, 100 = perfect). Jev's numbers are its official entry on the live board
(jev-1.13.0); ours are from the same kit and suite.
❯ Against its base model
Same prompt, same serving path, Qwen3.5-35B-A3B zero-shot vs JPT-35B-A3B: fine-tuning lifts every row but ANLI r1, where the strong base is level, and buys +15 points on typed decisions at the cost of calibration there (ECE 0.031 → 0.171).
Table, with calibration
| Benchmark (version, n) | What it tests | JPT-35B-A3B | Qwen3.5-35B-A3B zero-shot | Jev 1.13.0 |
|---|---|---|---|---|
| JevBench v1.4.0 public hard tier [2] (111) | hardest general decisions | 0.784 (ECE 0.066, Brier 0.297) | 0.703 (ECE 0.069, Brier 0.390) | — |
| Typed decisions test (ours, 2,000) | in-distribution typed decisions | 0.803 (ECE 0.171, Brier 0.323) | 0.651 (ECE 0.031, Brier 0.449) | — |
| ANLI r1 / r2 / r3 [4] (dev) | adversarial NLI | 0.773 / 0.680 / 0.737 | 0.777 / 0.650 / 0.703 | — |
| Banking77 [5] / MASSIVE 1.1 [6] (en / de / zh) | intent classification | 0.777 / 0.893 / 0.867 / 0.833 | 0.720 / 0.840 / 0.787 / 0.773 | — |
| AG News [7] / Emotion [8] / SST-5 [9] (test) | out-of-distribution classification | 0.917 / 0.593 / 0.593 | 0.887 / 0.517 / 0.543 | — |
| ScreenSpot-v2 [11] / Screen2Words [12] / ERQA [13] (images) | GUI grounding, screen summary, embodied reasoning | — | — | — |
| EnvBench v0.1 (ours) public / held-out [10] (skill 0–100) | sequential decisions in game envs | — | — | — |
Banking77, MASSIVE and typed rows are in-distribution (train splits in the mix, test items not). The image evals and EnvBench have not been run on this checkpoint yet.
Jev 1.13.0 has no official score on these splits (our own test/dev cuts, EnvBench, and the image sets), so its column is "—"; its official scores on the Decision Index versions of ANLI and BANKING77 are in the per-benchmark table above. Running the Jev API on these splits would fill them.
❯❯ Quick Start
Two pieces: an engine that holds the weights, and llm2jev (>= 0.6.1) in front of it, reading option probabilities off the engine. The bf16 weights are 66 GB, so below 80 GB per GPU use tensor parallelism.
⚡ SGLang (recommended)
python -m sglang.launch_server --model-path kirp/jpt-35b-a3b --port 30000 --tp 4 --mem-fraction-static 0.7 \
--context-length 32768 --mamba-scheduler-strategy extra_buffer & # Qwen3.5's DeltaNet layers need this flag
llm2jev --model kirp/jpt-35b-a3b --backend sglang --url http://127.0.0.1:30000 --port 8080 --temperature 1.087
Tested with SGLang 0.5.18 on 4 × A100 40 GB. --mem-fraction-static 0.7 leaves room for the vision tower's warm-up;
at the 0.8 we use for smaller sizes, the warm-up runs out of memory at TP 4.
🔁 vLLM
vllm serve kirp/jpt-35b-a3b --tensor-parallel-size 4 --max-logprobs 256 --return-tokens-as-token-ids \
--enable-scale-out --port 8000
llm2jev --model kirp/jpt-35b-a3b --backend vllm --url http://127.0.0.1:8000 --port 8080 --temperature 1.087
The three flags after --tensor-parallel-size are required: without them every request is a bare HTTP 400. Not yet
tested with this checkpoint.
📨 Ask it a question
import requests
r = requests.post("http://127.0.0.1:8080/v1/systemone", json={
"state": "Refund policy: full refund within 30 days of purchase; 50% until day 60; none after.\n"
"Order 1182 was bought on 3 March and returned on 20 April.",
"questions": {
"refund": {"type": "choice", "instructions": "What refund does order 1182 get?",
"criteria": {"full": "Full refund", "half": "50% refund", "none": "No refund"}},
"late": {"type": "noul", "instructions": "Was the return made after day 30?",
"criteria": {"true": "Yes", "false": "No"}}}})
print(r.json()["answers"]) # each answer has the per-option probabilities
🖼️ With images
state = [{"role": "user", "content": [
{"type": "image", "image": "https://example.com/screen.png"},
{"type": "text", "text": "Task: open the settings page. Numbered boxes mark clickable elements."}]}]
questions = {"click": {"type": "choice", "instructions": "Which box should be clicked?",
"criteria": {"1": None, "2": None, "3": None, "4": None, "5": None}}}
❯❯ Training
LoRA with a Brier loss on 49,221 typed questions, one epoch, trained with Megatron-SWIFT and merged into full weights.
| Part | What it is |
|---|---|
| Method | LoRA r=16, alpha 32 on every attention, DeltaNet, shared-expert and routed-expert projection of the language model; router and vision tower untouched |
| Loss | multi-class Brier over the option labels, on llm2jev's chat prompt with thinking disabled |
| Data | 49,221 questions in 32,835 records; one epoch over two option-shuffled copies (94,172 rows) |
| Schedule | lr 5e-5, cosine decay, 5% warm-up, global batch 64, 1,471 steps; final validation Brier loss 0.209 |
| Hardware | Megatron-SWIFT on 8 × A100 40 GB: expert parallelism 8, sequence parallelism, 2 h 17 min |
| Held out | no item from JevBench, EnvBench held-out seeds, the Decision Index frozen suite or our typed test split; 8-gram overlap check vs JevBench |
Data sources
The same mix as JPT-4B and JPT-9B (mix_train_env_v11):
- public classification, NLI, QA, preference and safety datasets;
- long legal and contract documents (ContractNLI, MAUD, LegalBench, ConditionalQA, ShARC);
- table and numeric reasoning (TAT-QA, MultiHiertt);
- multi-hop QA (MuSiQue, BEIR);
- agent and tool traces (Mind2Web, AgentTraj, ToolACE);
- community typed-decision sets;
- oracle-labelled rollouts from 20 small game and puzzle environments;
- programmatically generated rule-arithmetic items (dates, time zones, day counts, caps; labels computed by code);
- 1,289 long policy / contract / regulation documents (4,471 questions) written by an LLM (GPT-6 Luna) with no JevBench item shown to it.
❯❯ Limitations
- Arithmetic and dates are its weakest area: it answers from the evidence given and has no reasoning phase by design.
- Overconfident in-distribution. At the global temperature, typed-decisions test ECE is 0.171: confidence on in-distribution typed questions runs above accuracy.
- Up to 255 options are accepted; training covered up to 77 (Banking77), so beyond that quality is not established.
- English first. Other languages come only from a few multilingual classification sets.
- Images are untested on this checkpoint; they go zero-shot through the base vision tower.
- Heavy to host. 66 GB of bf16 weights need several GPUs, even though only 3B parameters are active per token.
❯❯ References
- TypeSafe AI. Jev. https://typesafe.ai
- F. Standhartinger. JevBench, v1.4.0. https://github.com/fstandhartinger/jevbench
- Decision Index, edition 0.2.1. https://huggingface.co/spaces/multimodalart/jev-decision-index
- Nie et al. Adversarial NLI. ACL 2020.
- Casanueva et al. Efficient Intent Detection with Dual Sentence Encoders (Banking77). NLP4ConvAI 2020.
- FitzGerald et al. MASSIVE. ACL 2023.
- Zhang et al. Character-level Convolutional Networks for Text Classification (AG News). NeurIPS 2015.
- Saravia et al. CARER: Contextualized Affect Representations for Emotion Recognition. EMNLP 2018.
- Socher et al. Recursive Deep Models for Semantic Compositionality (SST). EMNLP 2013.
- EnvBench, v0.1 (ours, frozen 2026-09-23; not yet public): programmatically solved game, planning and rule decisions with exact gold answers.
- Wu et al. OS-Atlas (ScreenSpot-v2). 2024.
- Wang et al. Screen2Words. UIST 2021.
- Gemini Robotics Team. Gemini Robotics (ERQA). 2025.
❯❯ License
CC BY-NC 4.0. The weights derive from Qwen3.5-35B-A3B (Apache-2.0), but some training datasets allow only non-commercial or research use, so the model is released for non-commercial use.
JPT-35B-A3B is an independent open model that implements a typed-decision interface (noul, choice and score questions answered with probabilities). It is not affiliated with, endorsed by or derived from TypeSafe AI or its Jev model, and it was not trained on Jev outputs.
- Downloads last month
- 21






