Jevling-2B-v0.1

Jevling-2B-v0.1 is a small System One decision model in the family of TypeSafe's Jev: you give it a state (any text — a transcript, a ticket, a document) and one or more typed questions (choice / yes-no / score), and it answers all of them in one forward pass, with no text generation, each as a calibrated probability distribution over the options. It is fine-tuned from google/gemma-4-E2B-it for on-device use (16 GB RAM), with special attention to Traditional-Chinese speech transcripts (kiosk ordering).

Quick start (transformers)

import torch
from transformers import AutoTokenizer, AutoModelForCausalLM

MODEL = "BricksDisplay/jevling-2b-v0.1"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, dtype=torch.bfloat16).to("cuda").eval()
LETTERS = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"
LETTER_IDS = [tok.encode(c, add_special_tokens=False)[0] for c in LETTERS]

def ask(state, questions):
    """questions: list of dicts {kind: 'choice'|'noul'|'score', text, options, descs (optional)}.
    noul options are always ['no','yes']. Returns one probability list per question (one forward pass)."""
    ids = ([tok.bos_token_id] if tok.bos_token_id is not None else []) + tok.encode(f"<state>\n{state}\n</state>\n", add_special_tokens=False)
    slots, sizes = [], []
    many = len(questions) > 1
    for k, q in enumerate(questions, 1):
        opts = ["no", "yes"] if q["kind"] == "noul" else q["options"]
        tag = {"noul": "yes/no", "score": "score"}.get(q["kind"], "choice")
        text = f"\nQuestion{' '+str(k) if many else ''} ({tag}): {q['text'].strip()}"
        text += "\nLevels:" if q["kind"] == "score" else ("\nOptions:" if q["kind"] == "choice" else "")
        for j, o in enumerate(opts):
            d = (q.get("descs") or [None] * len(opts))[j]
            text += f"\n({LETTERS[j]}) {o}" + (f" — {d}" if d else "")
        text += f"\nAnswer{' '+str(k) if many else ''}: ("
        ids += tok.encode(text, add_special_tokens=False)
        slots.append(len(ids) - 1); sizes.append(len(opts))
    with torch.no_grad():
        logits = model(input_ids=torch.tensor([ids], device=model.device)).logits[0]        # [T, vocab]
    return [torch.softmax(logits[s, LETTER_IDS[:n]].float(), 0).tolist() for s, n in zip(slots, sizes)]

state = "Customer: I was charged twice for the same subscription this month. Please refund the duplicate charge."
qs = [
  {"kind": "choice", "text": "Which team should handle this ticket?", "options": ["billing", "technical", "account"],
   "descs": ["payments and refunds", "a product fault", "login or profile settings"]},
  {"kind": "noul", "text": "Is the customer asking for a refund?"},
  {"kind": "score", "text": "How urgent is this?", "options": ["routine", "soon", "urgent", "critical"]},
]
for q, p in zip(qs, ask(state, qs)):
    print(q["text"], [round(x, 3) for x in p])

Output for that request:

Which team should handle this ticket? [0.998, 0.002, 0.0]     # billing
Is the customer asking for a refund? [0.001, 0.999]            # P(yes) = 1.00
How urgent is this? [0.374, 0.33, 0.2, 0.095]                  # expected level ≈ 1.0 of 0..3

Rules of the format: yes/no questions always use the options no, yes; score questions list ordered levels; give option descriptions whenever you have them; ask several questions per state — each is one extra answer slot, not a new prompt. The prompt layout above is the one the model was trained on; the same template is embedded in the GGUF as the named chat template system_one.

On device

Use the GGUF repo BricksDisplay/jevling-2b-v0.1-GGUF with the maintained llama.cpp implementation (tools/system-one on mybigday/system-one-llama.cpp, branch feat/system-one). Stock llama.cpp can load the weights but has no way to ask a typed question or read the answer.

Evaluation

All numbers are accuracy on datasets the models were not trained on, asked in the System One format (all questions of an item in one request; option descriptions given where the dataset has them). Items per dataset: 120–150 unless noted. Both models of the series are shown; this card's model in bold.

benchmark task Jevling-0.8B-v0.1 Jevling-2B-v0.1
MASSIVE (en-US) scenario classification, 18-way .675 .733
BBC News topic, 5-way .933 .958
TREC question type, 6-way .858 .850
PAWS paraphrase yes/no .508 .625
CommitmentBank NLI, 3-way .893 .875
StrategyQA yes/no reasoning .483 .567
PubMedQA yes/no/maybe .758 .667
SciQ 4-way science QA .942 .975
Social IQa 3-way .575 .725
TruthfulQA (MC) multiple choice .450 .633
XStoryCloze (en) 2-way .933 .958
QuALITY long-document 4-way QA .417 .500
RewardBench pairwise preference .600 .817
Hermes function-calling tool choice .996 .988
Financial PhraseBank sentiment, 3-way .608 .658
JevBench easy / original / hard (231 items) typed decisions 1.000 / .833 / .441 1.000 / .903 / .441
zh-TW kiosk set (ours, synthetic-derived, 255 states) intent acc / completeness AUROC / is-order / noise / size .969 / .971 / .996 / 1.000 / 1.000 .973 / .989 / .995 / 1.000 / 1.000

JevBench hard (.44) is where every open model we know of sits (TypeSafe's Jev API scores .73 on it); on the zh-TW kiosk set the same harness scores the two models at .971 / .973 and the Jev API at .931 — the kiosk set is our own synthetic-derived data, so read that comparison as "fit for the distribution it was built for", not as a general claim.

Limitations

  • Long-document, multi-hop, probability and date/time arithmetic questions are weak (JevBench hard .44; QuALITY .4–.5); compute arithmetic in code and put the result in the state.
  • Chinese coverage comes from synthetic kiosk-style data; validate on your own transcripts before relying on it.
  • When a question is undecidable the model picks the majority-style option; add an explicit "none of the above" option if you need abstention.
  • Not a chat model: it does not generate text.

Training data

Fine-tuned on a mix of public classification / QA / preference / tool-use datasets and synthetic Traditional-Chinese kiosk transcripts. None of the evaluation sets above were used for training.

Licence and release status

v0.1 is a research / non-commercial release. Some of the public datasets in the training mix carry non-commercial or research-only terms, so these weights are released under CC-BY-NC-4.0 on top of the Gemma Terms of Use (gemma-4 base).

Downloads last month
-
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BricksDisplay/jevling-2b-v0.1

Finetuned
(371)
this model
Quantizations
1 model