Jevling-E2B-v1

Jevling-E2B-v1 is a small System One decision model in the family of TypeSafe's Jev: you give it a state (any text — a transcript, a ticket, a document) and one or more typed questions (choice / yes-no / score), and it answers all of them in one forward pass, with no text generation, each as a calibrated probability distribution over the options. It is fine-tuned from google/gemma-4-E2B-it for on-device use (16 GB RAM), with special attention to Traditional-Chinese speech transcripts (kiosk ordering).

Quick start (transformers)

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

MODEL = "BricksDisplay/jevling-e2b-v1"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, dtype=torch.bfloat16).to("cuda").eval()   # on ROCm add attn_implementation="eager"
LETTERS = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"
LETTER_IDS = [tok.encode(c, add_special_tokens=False)[0] for c in LETTERS]

def ask(state, questions):
    """questions: list of dicts {kind: 'choice'|'noul'|'score', text, options, descs (optional)}.
    noul options are always ['no','yes']. Returns one probability list per question (one forward pass)."""
    ids = ([tok.bos_token_id] if tok.bos_token_id is not None else []) + tok.encode(f"<state>\n{state}\n</state>\n", add_special_tokens=False)
    slots, sizes = [], []
    many = len(questions) > 1
    for k, q in enumerate(questions, 1):
        opts = ["no", "yes"] if q["kind"] == "noul" else q["options"]
        tag = {"noul": "yes/no", "score": "score"}.get(q["kind"], "choice")
        text = f"\nQuestion{' '+str(k) if many else ''} ({tag}): {q['text'].strip()}"
        text += "\nLevels:" if q["kind"] == "score" else ("\nOptions:" if q["kind"] == "choice" else "")
        for j, o in enumerate(opts):
            d = (q.get("descs") or [None] * len(opts))[j]
            text += f"\n({LETTERS[j]}) {o}" + (f" — {d}" if d else "")
        text += f"\nAnswer{' '+str(k) if many else ''}: ("
        ids += tok.encode(text, add_special_tokens=False)
        slots.append(len(ids) - 1); sizes.append(len(opts))
    with torch.no_grad():
        logits = model(input_ids=torch.tensor([ids], device=model.device)).logits[0]        # [T, vocab]
    return [torch.softmax(logits[s, LETTER_IDS[:n]].float(), 0).tolist() for s, n in zip(slots, sizes)]

state = "Customer: I was charged twice for the same subscription this month. Please refund the duplicate charge."
qs = [
  {"kind": "choice", "text": "Which team should handle this ticket?", "options": ["billing", "technical", "account"],
   "descs": ["payments and refunds", "a product fault", "login or profile settings"]},
  {"kind": "noul", "text": "Is the customer asking for a refund?"},
  {"kind": "score", "text": "How urgent is this?", "options": ["routine", "soon", "urgent", "critical"]},
]
for q, p in zip(qs, ask(state, qs)):
    print(q["text"], [round(x, 3) for x in p])

Output for that request:

Which team should handle this ticket? [1.0, 0.0, 0.0]         # billing
Is the customer asking for a refund? [0.0, 1.0]                # P(yes) = 1.00
How urgent is this? [0.247, 0.671, 0.08, 0.001]                # expected level ≈ 0.8 of 0..3

Rules of the format: yes/no questions always use the options no, yes; score questions list ordered levels; give option descriptions whenever you have them; ask several questions per state — each is one extra answer slot, not a new prompt. The prompt layout above is the one the model was trained on; the same template is embedded in the GGUF as the named chat template system_one.

On device

Use the GGUF repo BricksDisplay/jevling-e2b-v1-GGUF with the maintained llama.cpp implementation (tools/system-one on mybigday/system-one-llama.cpp, branch feat/system-one). Stock llama.cpp can load the weights but has no way to ask a typed question or read the answer.

Evaluation

All numbers are accuracy on datasets the model was not trained on, asked in the System One format (all questions of an item in one request; option descriptions given where the dataset has them). Items per dataset: 120–150 unless noted.

benchmark task Jevling-0.8B-v1 Jevling-E2B-v1
MASSIVE (en-US) scenario classification, 18-way 0.667 0.742
BBC News topic, 5-way 0.917 0.967
TREC question type, 6-way 0.758 0.792
PAWS paraphrase yes/no 0.717 0.700
CommitmentBank NLI, 3-way 0.625 0.804
StrategyQA yes/no reasoning 0.500 0.525
PubMedQA yes/no/maybe 0.717 0.600
SciQ 4-way science QA 0.950 0.967
Social IQa 3-way 0.633 0.717
TruthfulQA (MC) multiple choice 0.467 0.642
XStoryCloze (en) 2-way 0.917 0.967
QuALITY long-document 4-way QA 0.425 0.567
RewardBench pairwise preference 0.567 0.733
Hermes function-calling tool choice 0.971 0.963
Financial PhraseBank sentiment, 3-way 0.658 0.667
JevBench easy / original / hard (231 items) typed decisions 1.000 / 0.861 / 0.450 1.000 / 0.944 / 0.432
zh-TW kiosk set (ours, synthetic-derived, 255 states) intent acc / completeness AUROC / is-order / noise / size 0.961 / 0.975 / 0.984 / 0.992 / 1.000 0.980 / 0.989 / 0.984 / 0.992 / 1.000

JevBench hard (.45) is where every open model we know of sits (TypeSafe's Jev API scores .73 on it); the zh-TW kiosk set is our own synthetic-derived data, so read that row as "fit for the distribution it was built for", not as a general claim.

Limitations

  • Long-document, multi-hop, probability and date/time arithmetic questions are weak (JevBench hard .44; QuALITY .4–.6); compute arithmetic in code and put the result in the state.
  • Chinese coverage comes from synthetic kiosk-style data; validate on your own transcripts before relying on it.
  • When a question is undecidable the model picks the majority-style option; add an explicit "none of the above" option if you need abstention.
  • Chat still works: the fine-tune only trains the answer slot, and spot checks show the base model's chat replies are essentially unchanged (chat quality was not benchmarked). The calibration temperature (T = 1.10) is folded into the final norm, so sampled chat output is slightly flatter than the base model's at the same sampling temperature; greedy decoding is unaffected.

Training data

Fine-tuned on commercially licensed public classification / QA / preference / tool-use / safety datasets and synthetic Traditional-Chinese kiosk transcripts. None of the evaluation sets above were used for training.

Licence

Apache-2.0 (same as the Gemma-4 base). Trained only on commercially usable data: public datasets under MIT, Apache-2.0, CC-BY-4.0, CC-BY-2.0, CC0, ODC-BY and CDLA-Sharing licences, plus our own synthetic data. CC-BY / ODC-BY sources require attribution; the per-dataset list is available on request.

Downloads last month
-
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BricksDisplay/jevling-e2b-v1

Finetuned
(384)
this model
Quantizations
1 model