ezjev-4b-s2

Project page: xnu.app/ezjev โ€” overview, downloads and a step-by-step quickstart. Code: github.com/everettjf/ezjev (MIT). A newer checkpoint, ezjev-4b-s3, is recommended for new use; this one is kept for reproducing the Decision Index run.

A typed-decision model (Jev-style /v1/systemone: choice, noul and score questions answered with a probability per option from one forward pass). LoRA fine-tune of Qwen/Qwen3.5-4B, merged into full weights, in two stages:

  1. Stage 1 (ezjev-4b): LoRA r=16 on all language-model linear layers (attention, Gated-DeltaNet, MLP), LR 1e-4, one epoch over ~85k rows / ~106k questions. Loss: cross-entropy + Brier over the option labels, llm2jev chat prompt (thinking off).
  2. Stage 2: a second LoRA at LR 5e-5 on ~30k rows targeting weak task families (RAG hallucination, product relevance, code-output selection, select-all-that-apply, clinical NLI, phishing emails, claim verification) with 40% replay.

Temperature 1.26, fitted by NLL on a held-out dev split.

Decision Index 0.3 (official)

Listed on the Jev Decision Index 0.3 board (2026-10-06): Full score 46.95, #35 of 111, the best of all models โ‰ค 5B (every entry ranked above it has at least 9B parameters).

Part (weight) Score
Public benchmarks, 37 (20%) 50.82
Private tests of the same skills (50%) 46.22
Private tasks from new domains (30%) 40.73
Full score 46.95

For comparison: Jev 1.13.0 scores 60.11 (#3); the next-best 4B entry, vLLM-SR Decision 2.0 Nox 4B, scores 44.95 (#38). Skill scores, 0 = chance. Source: the leaderboard's data/index.json.

Decision Index 0.2.1 (our own run)

51.15: a complete run of the frozen suite, with all 150,759 requests answered and a median latency of 35.5 ms on one RTX PRO 6000. Area skill: knowledge 33.5, language 60.2, retrieval 56.3, tools 69.9, arts 28.8. Run: everettjf/decision-index-results-ezjev-4b-s2.

Serving

pip install "vllm==0.30.0" "llm2jev>=0.6.1"
VLLM_USE_FLASHINFER_SAMPLER=0 vllm serve everettjf/ezjev-4b-s2 --port 8000 --max-logprobs 256 --return-tokens-as-token-ids \
    --max-model-len 131072 --gpu-memory-utilization 0.90 --additional-config '{"gdn_prefill_backend": "triton"}'
llm2jev --model everettjf/ezjev-4b-s2 --backend vllm --url http://127.0.0.1:8000 --port 8080 --temperature 1.26

The VLLM_USE_FLASHINFER_SAMPLER=0 and Triton GDN-prefill settings were used only because the job image had no CUDA toolkit for FlashInfer JIT. They do not change the logprobs llm2jev reads, and you can drop them where nvcc is available.

Training data

All data comes from train or dev splits of public datasets, plus code-generated items whose labels are computed by the generating program. No Decision Index suite rows were used. Every training row was also checked against the suite (exact-segment and 13-gram overlap) and dropped on any match. Sources: MNLI, SNLI, ANLI, BoolQ, BANKING77, CLINC150, MASSIVE, AG News, DBpedia, Yahoo Answers, Emotion, SST-5, TweetEval irony, WinoGrande, HellaSwag, MMLU auxiliary_train, ARC, CommonsenseQA, OpenBookQA, QASC, SciQ, AQuA-RAT, GSM8K, HotpotQA, Glaive function calling v2, SHP, HH-RLHF, ContractNLI, VAST, NLI4CT (SemEval-2024 Task 2), RAGTruth, iSarcasmEval, ACOS, ShARC, Amazon ESCI, GLUE QNLI, SGD (DSTC8), ToolACE, When2Call, Humicroedit, the New Yorker caption contest, and ealvaradob/phishing-dataset URLs.

Code: github.com/everettjf/ezjev.

Licences: some sources are non-commercial (for example ANLI, CC BY-NC 4.0), and some have no stated licence (VAST, NLI4CT, ACOS, Humicroedit). Check the licences before any commercial use.

Downloads last month
108
Safetensors
Model size
5B params
Tensor type
BF16
ยท
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for everettjf/ezjev-4b-s2

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(901)
this model
Finetunes
1 model
Quantizations
1 model

Space using everettjf/ezjev-4b-s2 1