Jevling-E2B-v0.1 β€” GGUF

Quantised build of BricksDisplay/jevling-e2b-v0.1 for on-device use (16 GB RAM). This is a System One decision model: one forward pass, no generated text, several typed questions per call, each answered as a calibrated probability. It does not work with stock llama.cpp chat/completion endpoints β€” they can load the weights but cannot ask a typed question or read the answer slot.

Use the maintained implementation

tools/system-one in mybigday/system-one-llama.cpp, branch feat/system-one:

git clone -b feat/system-one https://github.com/mybigday/system-one-llama.cpp
cd system-one-llama.cpp && cmake -B build && cmake --build build --target llama-system-one llama-system-one-batch -j

One call, three typed questions:

build/bin/llama-system-one -m jevling-e2b-v0.1-q8_0-embf16.gguf \
    --state "Customer: I was charged twice for the same subscription this month. Please refund the duplicate charge." \
    --choice "queue:Which team should handle this ticket?:billing=payments and refunds,technical=a product fault,account=login or profile settings" \
    --noul   "refund:Is the customer asking for a refund?" \
    --score  "urgency:How urgent is this?:routine,soon,urgent,critical" --json

Answers come back as probabilities per option (--json is the same shape the server returns: noul P(true), choice with a distribution, score as the expected level). Pass option descriptions β€” the model was trained with them. Use --system-one-request request.json for anything with colons or many questions; llama-system-one-batch for throughput. The prompt template and readout live inside the GGUF (tokenizer.chat_template.system_one; the readout is derived from the model, not declared); do not supply a chat template of your own.

Measured: β‰ˆ1.7 s per 5-question request on 16 CPU threads, β‰ˆ2.9 GB RSS; run with --swa-full. Keep flash-attention off on CPU and threads = physical cores.

Or the /v1/systemone endpoint

This file also answers upstream llama.cpp's decision-model endpoint, which needs a build that tokenizes the prompt piece by piece β€” the boundaries this model was trained on. That is one commit on top of upstream master:

git clone -b system-one/decision-segments https://github.com/mybigday/system-one-llama.cpp
cd system-one-llama.cpp && cmake -B build -DCMAKE_BUILD_TYPE=Release && cmake --build build --target llama-server -j
build/bin/llama-server -m jevling-e2b-v0.1-q8_0-embf16.gguf -c 2048 -t $(nproc) --n-outputs-max-per-seq 16
curl http://localhost:8080/v1/systemone -H "Content-Type: application/json" -d '{
  "state": "Customer: I was charged twice and nobody replied.",
  "questions": {
    "queue":   {"type": "choice", "instructions": "Which team should handle this?",
                "criteria": {"billing": "payments and refunds", "technical": "a product fault"}},
    "refund":  {"type": "noul",   "instructions": "Is the customer asking for a refund?"},
    "urgency": {"type": "score",  "instructions": "How urgent is this?",
                "criteria": ["routine", "soon", "urgent", "critical"]}
  }
}'

Stock llama.cpp must not be used for this. It will load the file and answer, and its answers will be wrong without saying so: a single pass over the whole prompt merges across one of the training boundaries and comes out one token shorter, which moves a probability by up to 0.28 and changes a few decisions in a thousand. Nothing is raised, because the separator the template writes is simply undefined there and renders as empty. The branch above is a requirement, not a suggestion.

Measured against the fp32 reference these weights were validated on, 255 states / 1275 answers (every published precision is in the file table below): the jevling-e2b-v0.1-q8_0-embf16.gguf build's decision agreement is noul 1.0000 Β· choice 1.0000 Β· score 1.0000, accuracy unchanged on all three, 2 of 1275 flips (both ties).

Files

file size decision agreement (noul / choice / score) decisions moved / 1275
jevling-e2b-v0.1-BF16.gguf 9.30 GB the reference 0 original trained precision
jevling-e2b-v0.1-q8_0-embf16.gguf 7.55 GB 1.0000 / 1.0000 / 1.0000 2 recommended
jevling-e2b-v0.1-Q8_0.gguf 4.97 GB 0.9987 / 1.0000 / 1.0000 1 fine, more RAM
jevling-e2b-v0.1-Q4_K_M.gguf 3.44 GB 0.9934 / 1.0000 / 1.0000 7 not recommended
mmproj-jevling-e2b-v0.1-BF16.gguf 987 MB β€” β€” image/audio projector
mmproj-jevling-e2b-v0.1-Q8_0.gguf 557 MB β€” β€” image/audio projector

Agreement is measured on a stock build over /v1/systemone against the fp32 reference these weights were validated on, 255 states / 1275 answers. The bar is the one this model ships under: β‰₯ 0.99 non-boundary argmax agreement per question kind, with accuracy dropping no more than 0.01 β€” a decision model is judged on whether it decides the same thing, not on probability error.

Q4_K_M is not recommended, which is unusual enough to be worth stating plainly: take Q8_0 even when space is tight. It clears the bar here by only 0.4 pp, and the same bytes read 0.9895 on an earlier host, so the model sits on the line rather than above it β€” the other three models in this series miss it outright. The damage lands on noul questions sitting near their threshold; choice and score are unaffected.

On E2B, prefer q8_0-embf16 even though the file is bigger. Quantising the per-layer token embeddings raises resident memory instead of lowering it β€” measured on this architecture, f16 embeddings 2.9 GB RSS against Q8_0's 5.3 GB β€” so the larger file is the lighter process.

Evaluation

All numbers are accuracy on datasets the models were not trained on, asked in the System One format (all questions of an item in one request; option descriptions given where the dataset has them). Items per dataset: 120–150 unless noted. Both models of the series are shown; this card's model in bold.

benchmark task Jevling-0.8B-v0.1 Jevling-E2B-v0.1
MASSIVE (en-US) scenario classification, 18-way .675 .733
BBC News topic, 5-way .933 .958
TREC question type, 6-way .858 .850
PAWS paraphrase yes/no .508 .625
CommitmentBank NLI, 3-way .893 .875
StrategyQA yes/no reasoning .483 .567
PubMedQA yes/no/maybe .758 .667
SciQ 4-way science QA .942 .975
Social IQa 3-way .575 .725
TruthfulQA (MC) multiple choice .450 .633
XStoryCloze (en) 2-way .933 .958
QuALITY long-document 4-way QA .417 .500
RewardBench pairwise preference .600 .817
Hermes function-calling tool choice .996 .988
Financial PhraseBank sentiment, 3-way .608 .658
JevBench easy / original / hard (231 items) typed decisions 1.000 / .833 / .441 1.000 / .903 / .441
zh-TW kiosk set (ours, synthetic-derived, 255 states) intent acc / completeness AUROC / is-order / noise / size .969 / .971 / .996 / 1.000 / 1.000 .973 / .989 / .995 / 1.000 / 1.000

JevBench hard (.44) is where every open model we know of sits (TypeSafe's Jev API scores .73 on it); on the zh-TW kiosk set the same harness scores the two models at .971 / .973 and the Jev API at .931 β€” the kiosk set is our own synthetic-derived data, so read that comparison as "fit for the distribution it was built for", not as a general claim.

Limitations

  • Long-document, multi-hop, probability and date/time arithmetic questions are weak (JevBench hard .44; QuALITY .4–.5); compute arithmetic in code and put the result in the state.
  • Chinese coverage comes from synthetic kiosk-style data; validate on your own transcripts before relying on it.
  • When a question is undecidable the model picks the majority-style option; add an explicit "none of the above" option if you need abstention.
  • Not a chat model: it does not generate text.

Licence and release status

v0.1 is a research / non-commercial release. Some of the public datasets in the training mix carry non-commercial or research-only terms, so these weights are released under CC-BY-NC-4.0 on top of the Gemma Terms of Use (gemma-4 base).

Downloads last month
579
GGUF
Model size
5B params
Architecture
gemma4
Hardware compatibility
Log In to add your hardware

4-bit

8-bit

16-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for BricksDisplay/jevling-e2b-v0.1-GGUF

Quantized
(1)
this model

Collection including BricksDisplay/jevling-e2b-v0.1-GGUF