Instructions to use BricksDisplay/jevling-2b-v0.1 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use BricksDisplay/jevling-2b-v0.1 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="BricksDisplay/jevling-2b-v0.1")# Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("BricksDisplay/jevling-2b-v0.1") model = AutoModelForMultimodalLM.from_pretrained("BricksDisplay/jevling-2b-v0.1", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Jevling-2B-v0.1
Jevling-2B-v0.1 is a small System One decision model in the family of TypeSafe's Jev: you give it a state (any text — a transcript, a ticket, a document) and one or more typed questions (choice / yes-no / score), and it answers all of them in one forward pass, with no text generation, each as a calibrated probability distribution over the options. It is fine-tuned from google/gemma-4-E2B-it for on-device use (16 GB RAM), with special attention to Traditional-Chinese speech transcripts (kiosk ordering).
Quick start (transformers)
import torch
from transformers import AutoTokenizer, AutoModelForCausalLM
MODEL = "BricksDisplay/jevling-2b-v0.1"
tok = AutoTokenizer.from_pretrained(MODEL)
model = AutoModelForCausalLM.from_pretrained(MODEL, dtype=torch.bfloat16).to("cuda").eval()
LETTERS = "ABCDEFGHIJKLMNOPQRSTUVWXYZabcdefghijklmnopqrstuvwxyz"
LETTER_IDS = [tok.encode(c, add_special_tokens=False)[0] for c in LETTERS]
def ask(state, questions):
"""questions: list of dicts {kind: 'choice'|'noul'|'score', text, options, descs (optional)}.
noul options are always ['no','yes']. Returns one probability list per question (one forward pass)."""
ids = ([tok.bos_token_id] if tok.bos_token_id is not None else []) + tok.encode(f"<state>\n{state}\n</state>\n", add_special_tokens=False)
slots, sizes = [], []
many = len(questions) > 1
for k, q in enumerate(questions, 1):
opts = ["no", "yes"] if q["kind"] == "noul" else q["options"]
tag = {"noul": "yes/no", "score": "score"}.get(q["kind"], "choice")
text = f"\nQuestion{' '+str(k) if many else ''} ({tag}): {q['text'].strip()}"
text += "\nLevels:" if q["kind"] == "score" else ("\nOptions:" if q["kind"] == "choice" else "")
for j, o in enumerate(opts):
d = (q.get("descs") or [None] * len(opts))[j]
text += f"\n({LETTERS[j]}) {o}" + (f" — {d}" if d else "")
text += f"\nAnswer{' '+str(k) if many else ''}: ("
ids += tok.encode(text, add_special_tokens=False)
slots.append(len(ids) - 1); sizes.append(len(opts))
with torch.no_grad():
logits = model(input_ids=torch.tensor([ids], device=model.device)).logits[0] # [T, vocab]
return [torch.softmax(logits[s, LETTER_IDS[:n]].float(), 0).tolist() for s, n in zip(slots, sizes)]
state = "Customer: I was charged twice for the same subscription this month. Please refund the duplicate charge."
qs = [
{"kind": "choice", "text": "Which team should handle this ticket?", "options": ["billing", "technical", "account"],
"descs": ["payments and refunds", "a product fault", "login or profile settings"]},
{"kind": "noul", "text": "Is the customer asking for a refund?"},
{"kind": "score", "text": "How urgent is this?", "options": ["routine", "soon", "urgent", "critical"]},
]
for q, p in zip(qs, ask(state, qs)):
print(q["text"], [round(x, 3) for x in p])
Output for that request:
Which team should handle this ticket? [0.998, 0.002, 0.0] # billing
Is the customer asking for a refund? [0.001, 0.999] # P(yes) = 1.00
How urgent is this? [0.374, 0.33, 0.2, 0.095] # expected level ≈ 1.0 of 0..3
Rules of the format: yes/no questions always use the options no, yes; score questions list ordered levels; give option descriptions whenever you have them; ask several questions per state — each is one extra answer slot, not a new prompt. The prompt layout above is the one the model was trained on; the same template is embedded in the GGUF as the named chat template system_one.
On device
Use the GGUF repo BricksDisplay/jevling-2b-v0.1-GGUF with the maintained llama.cpp implementation (tools/system-one on mybigday/system-one-llama.cpp, branch feat/system-one). Stock llama.cpp can load the weights but has no way to ask a typed question or read the answer.
Evaluation
All numbers are accuracy on datasets the models were not trained on, asked in the System One format (all questions of an item in one request; option descriptions given where the dataset has them). Items per dataset: 120–150 unless noted. Both models of the series are shown; this card's model in bold.
| benchmark | task | Jevling-0.8B-v0.1 | Jevling-2B-v0.1 |
|---|---|---|---|
| MASSIVE (en-US) | scenario classification, 18-way | .675 | .733 |
| BBC News | topic, 5-way | .933 | .958 |
| TREC | question type, 6-way | .858 | .850 |
| PAWS | paraphrase yes/no | .508 | .625 |
| CommitmentBank | NLI, 3-way | .893 | .875 |
| StrategyQA | yes/no reasoning | .483 | .567 |
| PubMedQA | yes/no/maybe | .758 | .667 |
| SciQ | 4-way science QA | .942 | .975 |
| Social IQa | 3-way | .575 | .725 |
| TruthfulQA (MC) | multiple choice | .450 | .633 |
| XStoryCloze (en) | 2-way | .933 | .958 |
| QuALITY | long-document 4-way QA | .417 | .500 |
| RewardBench | pairwise preference | .600 | .817 |
| Hermes function-calling | tool choice | .996 | .988 |
| Financial PhraseBank | sentiment, 3-way | .608 | .658 |
| JevBench easy / original / hard (231 items) | typed decisions | 1.000 / .833 / .441 | 1.000 / .903 / .441 |
| zh-TW kiosk set (ours, synthetic-derived, 255 states) | intent acc / completeness AUROC / is-order / noise / size | .969 / .971 / .996 / 1.000 / 1.000 | .973 / .989 / .995 / 1.000 / 1.000 |
JevBench hard (.44) is where every open model we know of sits (TypeSafe's Jev API scores .73 on it); on the zh-TW kiosk set the same harness scores the two models at .971 / .973 and the Jev API at .931 — the kiosk set is our own synthetic-derived data, so read that comparison as "fit for the distribution it was built for", not as a general claim.
Limitations
- Long-document, multi-hop, probability and date/time arithmetic questions are weak (JevBench hard .44; QuALITY .4–.5); compute arithmetic in code and put the result in the state.
- Chinese coverage comes from synthetic kiosk-style data; validate on your own transcripts before relying on it.
- When a question is undecidable the model picks the majority-style option; add an explicit "none of the above" option if you need abstention.
- Not a chat model: it does not generate text.
Training data
Fine-tuned on a mix of public classification / QA / preference / tool-use datasets and synthetic Traditional-Chinese kiosk transcripts. None of the evaluation sets above were used for training.
Licence and release status
v0.1 is a research / non-commercial release. Some of the public datasets in the training mix carry non-commercial or research-only terms, so these weights are released under CC-BY-NC-4.0 on top of the Gemma Terms of Use (gemma-4 base).
- Downloads last month
- -