j33-jev

A small "System One" decision model from J33.AI. You give it a case (a ticket, an invoice, an alert, an agent trace) and a few fixed questions. It picks an answer to each question and gives a calibrated probability, so you can let it decide alone when it is sure and send the rest to a person.

It is Google's EmbeddingGemma 2 encoder plus a small scoring layer, fine-tuned on the train split of LocalLLaMA/typed-decisions. The design follows the open-source Laya decision model: one sequence per question, a <mask> marker in front of every option, a score read at each marker, and one fitted temperature per question type.

"Jev-style" describes the interface (typed questions in, calibrated probabilities out). This model is not affiliated with or endorsed by TypeSafe, the maker of Jev.

Results

Typed-decisions test set: 400 cases, 2,000 decisions. Scoring from J33's benchmark: +1 for a right decision made alone, −5 for a wrong one, 0 when passed to a person. The model decides alone only when its confidence clears the break-even point (83.3%).

Model Trained on this train split? Accuracy ECE Work saved per 100 decisions Decides alone
j33-jev (this model) Yes 75.2% 0.019 +22.5 39.9%
Laya, fitted, own temperatures Yes 76.9% 0.137 +17.1 17.7%
Jev (TypeSafe, hosted, as delivered) No 73.7% 0.033 +13.0 37.6%
Julia-1 Unknown (private data) 73.0% 0.229 −43.5 90.3%

Read this fairly:

  • This model and the fitted Laya were trained on the benchmark's own train split; Jev was used as delivered with no training on it. That is a head start.
  • Gold labels are the average of three samples from an AI labeller on synthetic cases. Its self-agreement is 73.5%, so part of any score above that is learning the labeller's habits.
  • One training run per variant, seed 0. Median latency was 33 ms per case on the training GPU (fp32), not on the hardware used for the other models.

Full write-up: Building J33-Jev: is your own AI decision model worth it?. Earlier comparison of ready-made models: System One AI models: which ones can you use in production?

Usage

# pip install torch transformers safetensors pillow huggingface_hub
from huggingface_hub import snapshot_download
import sys

path = snapshot_download("J33-AI/j33-jev")
sys.path.insert(0, path)
import gemma_jev

model = gemma_jev.load(f"{path}/model.safetensors", base="google/embeddinggemma-2",
                       vision=False, head="none", device="cpu")

state = {"task": "Rotate the expired TLS certificate on the staging load balancer.",
         "trace_summary": {"steps": 11, "tool_errors": 0, "constraint_violations": 0}}
questions = {
    "action": {
        "type": "choice",
        "instructions": "What should the observability system do with this trace?",
        "criteria": {
            "continue": "Let the agent proceed without interruption.",
            "observe": "Keep running, but flag the trace for later sampling.",
            "human_review": "Queue this trace for a human to review.",
            "stop": "Halt the agent now.",
        },
    },
}
probs = model.predict(state, questions)  # {"action": {"continue": 0.41, "observe": 0.55, ...}}

Question types: choice (pick one label), score (ordered levels, criteria is a list), noul (yes/no, criteria has false/true). State can be any JSON-serialisable value; it is truncated to 4,000 characters. base is only used for the encoder config and tokenizer; all weights come from model.safetensors.

To decide alone, take the top option when its probability is at least your break-even confidence (0.833 for +1/−5 scoring), otherwise send the case to a person.

Training

  • Base: google/embeddinggemma-2 text encoder (vision tower off), 271.6M parameters in total including a 0.59M scorer. No extra transformer layer on top (head="none"); a bigger GQA head scored worse in our runs.
  • Data: 1,080 train cases (5,400 decisions); 120 cases (600 decisions) held out for validation and temperature fitting. The test set was never used for fitting.
  • Soft cross-entropy against the soft gold labels, options shuffled each batch, AdamW (encoder LR 2.5e-5, head LR 1e-4, weight decay 0.01), OneCycle with 6% warm-up, effective batch 32, 4 epochs, seed 0.
  • Calibration: one temperature per question type, fitted with LBFGS to the held-out set's top labels. Fitted temperatures: choice 0.467, score 0.455, noul 0.663.

Limitations

  • Trained and tested on synthetic cases from four workflows (customer service, invoice processing, security incidents, agent-trace observability). Fine-tune and re-fit the temperatures on your own decisions before relying on it elsewhere.
  • Text only. The training code supports images, but this checkpoint was trained without them.
  • English only.

Licence

Apache 2.0, the same as the base model and the training data as listed on Hugging Face.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for J33-AI/j33-jev

Finetuned
(44)
this model

Dataset used to train J33-AI/j33-jev

Evaluation results