j33-jev
A small "System One" decision model from J33.AI. You give it a case (a ticket, an invoice, an alert, an agent trace) and a few fixed questions. It picks an answer to each question and gives a calibrated probability, so you can let it decide alone when it is sure and send the rest to a person.
It is Google's EmbeddingGemma 2 encoder plus a small scoring layer, fine-tuned on the train split of LocalLLaMA/typed-decisions. The design follows the open-source Laya decision model: one sequence per question, a <mask> marker in front of every option, a score read at each marker, and one fitted temperature per question type.
"Jev-style" describes the interface (typed questions in, calibrated probabilities out). This model is not affiliated with or endorsed by TypeSafe, the maker of Jev.
Results
Typed-decisions test set: 400 cases, 2,000 decisions. Scoring from J33's benchmark: +1 for a right decision made alone, −5 for a wrong one, 0 when passed to a person. The model decides alone only when its confidence clears the break-even point (83.3%).
| Model | Trained on this train split? | Accuracy | ECE | Work saved per 100 decisions | Decides alone |
|---|---|---|---|---|---|
| j33-jev (this model) | Yes | 75.2% | 0.019 | +22.5 | 39.9% |
| Laya, fitted, own temperatures | Yes | 76.9% | 0.137 | +17.1 | 17.7% |
| Jev (TypeSafe, hosted, as delivered) | No | 73.7% | 0.033 | +13.0 | 37.6% |
| Julia-1 | Unknown (private data) | 73.0% | 0.229 | −43.5 | 90.3% |
Read this fairly:
- This model and the fitted Laya were trained on the benchmark's own train split; Jev was used as delivered with no training on it. That is a head start.
- Gold labels are the average of three samples from an AI labeller on synthetic cases. Its self-agreement is 73.5%, so part of any score above that is learning the labeller's habits.
- One training run per variant, seed 0. Median latency was 33 ms per case on the training GPU (fp32), not on the hardware used for the other models.
Full write-up: Building J33-Jev: is your own AI decision model worth it?. Earlier comparison of ready-made models: System One AI models: which ones can you use in production?
Usage
# pip install torch transformers safetensors pillow huggingface_hub
from huggingface_hub import snapshot_download
import sys
path = snapshot_download("J33-AI/j33-jev")
sys.path.insert(0, path)
import gemma_jev
model = gemma_jev.load(f"{path}/model.safetensors", base="google/embeddinggemma-2",
vision=False, head="none", device="cpu")
state = {"task": "Rotate the expired TLS certificate on the staging load balancer.",
"trace_summary": {"steps": 11, "tool_errors": 0, "constraint_violations": 0}}
questions = {
"action": {
"type": "choice",
"instructions": "What should the observability system do with this trace?",
"criteria": {
"continue": "Let the agent proceed without interruption.",
"observe": "Keep running, but flag the trace for later sampling.",
"human_review": "Queue this trace for a human to review.",
"stop": "Halt the agent now.",
},
},
}
probs = model.predict(state, questions) # {"action": {"continue": 0.41, "observe": 0.55, ...}}
Question types: choice (pick one label), score (ordered levels, criteria is a list), noul (yes/no, criteria has false/true). State can be any JSON-serialisable value; it is truncated to 4,000 characters. base is only used for the encoder config and tokenizer; all weights come from model.safetensors.
To decide alone, take the top option when its probability is at least your break-even confidence (0.833 for +1/−5 scoring), otherwise send the case to a person.
Training
- Base:
google/embeddinggemma-2text encoder (vision tower off), 271.6M parameters in total including a 0.59M scorer. No extra transformer layer on top (head="none"); a bigger GQA head scored worse in our runs. - Data: 1,080 train cases (5,400 decisions); 120 cases (600 decisions) held out for validation and temperature fitting. The test set was never used for fitting.
- Soft cross-entropy against the soft gold labels, options shuffled each batch, AdamW (encoder LR 2.5e-5, head LR 1e-4, weight decay 0.01), OneCycle with 6% warm-up, effective batch 32, 4 epochs, seed 0.
- Calibration: one temperature per question type, fitted with LBFGS to the held-out set's top labels. Fitted temperatures: choice 0.467, score 0.455, noul 0.663.
Limitations
- Trained and tested on synthetic cases from four workflows (customer service, invoice processing, security incidents, agent-trace observability). Fine-tune and re-fit the temperatures on your own decisions before relying on it elsewhere.
- Text only. The training code supports images, but this checkpoint was trained without them.
- English only.
Licence
Apache 2.0, the same as the base model and the training data as listed on Hugging Face.
Model tree for J33-AI/j33-jev
Base model
google/embeddinggemma-2Dataset used to train J33-AI/j33-jev
Evaluation results
- accuracy on LocalLLaMA/typed-decisions (test, 2,000 decisions)test set self-reported0.752
- ece on LocalLLaMA/typed-decisions (test, 2,000 decisions)test set self-reported0.019