Jevling-E2B-v1 β€” GGUF

Quantised build of BricksDisplay/jevling-e2b-v1 for on-device use (16 GB RAM). This is a System One decision model: one forward pass, no generated text, several typed questions per call, each answered as a calibrated probability. It does not work with stock llama.cpp chat/completion endpoints β€” they can load the weights but cannot ask a typed question or read the answer slot.

Use the maintained implementation

tools/system-one in mybigday/system-one-llama.cpp, branch feat/system-one:

git clone -b feat/system-one https://github.com/mybigday/system-one-llama.cpp
cd system-one-llama.cpp && cmake -B build && cmake --build build --target llama-system-one llama-system-one-batch -j

One call, three typed questions:

build/bin/llama-system-one -m jevling-e2b-v1-q8_0-embf16.gguf \
    --state "Customer: I was charged twice for the same subscription this month. Please refund the duplicate charge." \
    --choice "queue:Which team should handle this ticket?:billing=payments and refunds,technical=a product fault,account=login or profile settings" \
    --noul   "refund:Is the customer asking for a refund?" \
    --score  "urgency:How urgent is this?:routine,soon,urgent,critical" --json

Answers come back as probabilities per option (--json is the same shape the server returns: noul P(true), choice with a distribution, score as the expected level). Pass option descriptions β€” the model was trained with them. Use --system-one-request request.json for anything with colons or many questions; llama-system-one-batch for throughput. The prompt template and readout live inside the GGUF (tokenizer.chat_template.system_one; the readout is derived from the model, not declared); do not supply a chat template of your own.

Measured: β‰ˆ1.7 s per 5-question request on 16 CPU threads, β‰ˆ2.9 GB RSS; run with --swa-full. Keep flash-attention off on CPU and threads = physical cores.

Or the /v1/systemone endpoint

This file also answers upstream llama.cpp's decision-model endpoint, which needs a build that tokenizes the prompt piece by piece β€” the boundaries this model was trained on. That is one commit on top of upstream master:

git clone -b system-one/decision-segments https://github.com/mybigday/system-one-llama.cpp
cd system-one-llama.cpp && cmake -B build -DCMAKE_BUILD_TYPE=Release && cmake --build build --target llama-server -j
build/bin/llama-server -m jevling-e2b-v1-q8_0-embf16.gguf -c 2048 -t $(nproc) --n-outputs-max-per-seq 16
curl http://localhost:8080/v1/systemone -H "Content-Type: application/json" -d '{
  "state": "Customer: I was charged twice and nobody replied.",
  "questions": {
    "queue":   {"type": "choice", "instructions": "Which team should handle this?",
                "criteria": {"billing": "payments and refunds", "technical": "a product fault"}},
    "refund":  {"type": "noul",   "instructions": "Is the customer asking for a refund?"},
    "urgency": {"type": "score",  "instructions": "How urgent is this?",
                "criteria": ["routine", "soon", "urgent", "critical"]}
  }
}'

Stock llama.cpp must not be used for this. It will load the file and answer, and its answers will be wrong without saying so: a single pass over the whole prompt merges across one of the training boundaries and comes out one token shorter, which moves a probability by up to 0.28 and changes a few decisions in a thousand. Nothing is raised, because the separator the template writes is simply undefined there and renders as empty. The branch above is a requirement, not a suggestion.

Measured against the fp32 reference these weights were validated on, 255 states / 1275 answers: its f32 weights read 1.5e-03 worst-case probability deviation from that reference with 0 of 1275 argmax differences β€” the figure is larger than the 0.8B's because gemma-4 uses GELU and llama.cpp's CPU GELU goes through an fp16 table, which changes no decision and is upstream's own accepted behaviour, and this file's decision agreement is noul 1.0000 Β· choice 1.0000 Β· score 1.0000, accuracy unchanged on all three, 0 flips.

Evaluation

All numbers are accuracy on datasets the model was not trained on, asked in the System One format (all questions of an item in one request; option descriptions given where the dataset has them). Items per dataset: 120–150 unless noted.

benchmark task Jevling-0.8B-v1 Jevling-E2B-v1
MASSIVE (en-US) scenario classification, 18-way 0.667 0.742
BBC News topic, 5-way 0.917 0.967
TREC question type, 6-way 0.758 0.792
PAWS paraphrase yes/no 0.717 0.700
CommitmentBank NLI, 3-way 0.625 0.804
StrategyQA yes/no reasoning 0.500 0.525
PubMedQA yes/no/maybe 0.717 0.600
SciQ 4-way science QA 0.950 0.967
Social IQa 3-way 0.633 0.717
TruthfulQA (MC) multiple choice 0.467 0.642
XStoryCloze (en) 2-way 0.917 0.967
QuALITY long-document 4-way QA 0.425 0.567
RewardBench pairwise preference 0.567 0.733
Hermes function-calling tool choice 0.971 0.963
Financial PhraseBank sentiment, 3-way 0.658 0.667
JevBench easy / original / hard (231 items) typed decisions 1.000 / 0.861 / 0.450 1.000 / 0.944 / 0.432
zh-TW kiosk set (ours, synthetic-derived, 255 states) intent acc / completeness AUROC / is-order / noise / size 0.961 / 0.975 / 0.984 / 0.992 / 1.000 0.980 / 0.989 / 0.984 / 0.992 / 1.000

JevBench hard (.45) is where every open model we know of sits (TypeSafe's Jev API scores .73 on it); the zh-TW kiosk set is our own synthetic-derived data, so read that row as "fit for the distribution it was built for", not as a general claim.

Limitations

  • Long-document, multi-hop, probability and date/time arithmetic questions are weak (JevBench hard .44; QuALITY .4–.6); compute arithmetic in code and put the result in the state.
  • Chinese coverage comes from synthetic kiosk-style data; validate on your own transcripts before relying on it.
  • When a question is undecidable the model picks the majority-style option; add an explicit "none of the above" option if you need abstention.
  • Primarily a decision model, not a chat assistant. The fine-tune only trains the answer slot, so the base model's chat ability is retained (spot-checked in transformers, not benchmarked); this GGUF build is intended for System One use.

Licence

Apache-2.0 (same as the Gemma-4 base). Trained only on commercially usable data: public datasets under MIT, Apache-2.0, CC-BY-4.0, CC-BY-2.0, CC0, ODC-BY and CDLA-Sharing licences, plus our own synthetic data. CC-BY / ODC-BY sources require attribution; the per-dataset list is available on request.

Downloads last month
166
GGUF
Model size
5B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for BricksDisplay/jevling-e2b-v1-GGUF

Quantized
(1)
this model