JEV-9B-GGUF

autotrust/JEV-9B is AutoTrust AI's first integrated System 1 + System 2 open model, built on a frozen, bit-identical Qwen3.5-9B backbone using a "Blocks of Experts" recipe: System 2 is ordinary text generation through the untouched base lm_head (70.7% HumanEval pass@1, identical to base Qwen3.5-9B with all 164 completions byte-identical), while System 1 is a small, detachable 40.2M-parameter LoRA plus a 24-slot fp32 decision head that answers typed noul (yes/no), choice (2–16 options), or score (0–5 scale) questions in a single forward pass, distilled from the closed, hosted TypeSafe Jev 1.13's own output distributions via the Apache-2.0 SargeDev/jev-distill-corpus-v3 corpus. On 25,376 Jev-labelled held-out rows, JEV-9B reaches a mean KL divergence of just ≈0.019 nats from the teacher's distributions (essentially indistinguishable at that resolution, including reproducing several of the teacher's known mistakes), 90.2% choice top-1 agreement, 0.994 noul AUROC, and an ECE of 0.0007 with no post-hoc correction needed, while also generalizing to unseen task families (KL 0.234, top-1 91.8% on out-of-distribution Open-Jev rows) and reaching 90–97% of the teacher's accuracy on an independent third-party benchmark with human gold labels. It is dramatically faster than the hosted API — a single decision takes ~90ms median versus 238–301ms for the hosted service, and one B200 GPU sustains ~15x the throughput — and both systems are served from one set of weights via a single vLLM engine, with a request routed to either path per-call; its larger sibling, autotrust/JEV-27B, trades some of this speed for closer teacher fidelity, better OOD transfer, and stronger System 2 generation (78.0% HumanEval). The model, its LoRA adapter, decision head, and training/evaluation reports are all released under Apache-2.0, and AutoTrust AI states it is an independent, unaffiliated reproduction sharing no code or weights with TypeSafe AI.

Model Files

File Name Quant Type File Size File Link Description
JEV-9B.BF16.gguf BF16 17.9 GB Link Full BF16 weights. Highest quality, largest file size.
JEV-9B.Q3_K_L.gguf Q3_K_L 4.93 GB Link Lower quality but usable, good for low RAM availability.
JEV-9B.Q3_K_M.gguf Q3_K_M 4.62 GB Link Low quality.
JEV-9B.Q4_K_M.gguf Q4_K_M 5.63 GB Link Good quality, default size for most use cases, recommended.
JEV-9B.Q4_K_S.gguf Q4_K_S 5.35 GB Link Slightly lower quality with more space savings, recommended.
JEV-9B.Q5_K_M.gguf Q5_K_M 6.47 GB Link High quality, recommended.
JEV-9B.Q5_K_S.gguf Q5_K_S 6.31 GB Link High quality, recommended.
JEV-9B.Q6_K.gguf Q6_K 7.36 GB Link Very high quality, near perfect, recommended.

llama.cpp

LLM inference in C/C++ — https://github.com/ggml-org/llama.cpp

Downloads last month
-
GGUF
Model size
9B params
Architecture
qwen35
Hardware compatibility
Log In to add your hardware

3-bit

4-bit

5-bit

6-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for prithivMLmods/JEV-9B-GGUF

Finetuned
Qwen/Qwen3.5-9B
Finetuned
autotrust/JEV-9B
Quantized
(1)
this model

Collection including prithivMLmods/JEV-9B-GGUF