Instructions to use BricksDisplay/jevling-e2b-v1-GGUF with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- llama.cpp
How to use BricksDisplay/jevling-e2b-v1-GGUF with llama.cpp:
Install (macOS, Linux)
curl -LsSf https://llama.app/install.sh | sh # Start a local OpenAI-compatible server with a web UI: llama serve -hf BricksDisplay/jevling-e2b-v1-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf BricksDisplay/jevling-e2b-v1-GGUF:BF16
Install from WinGet (Windows)
winget install llama.cpp # Start a local OpenAI-compatible server with a web UI: llama serve -hf BricksDisplay/jevling-e2b-v1-GGUF:BF16 # Run inference directly in the terminal: llama cli -hf BricksDisplay/jevling-e2b-v1-GGUF:BF16
Use pre-built binary
# Download pre-built binary from: # https://github.com/ggerganov/llama.cpp/releases # Start a local OpenAI-compatible server with a web UI: ./llama-server -hf BricksDisplay/jevling-e2b-v1-GGUF:BF16 # Run inference directly in the terminal: ./llama-cli -hf BricksDisplay/jevling-e2b-v1-GGUF:BF16
Build from source code
git clone https://github.com/ggerganov/llama.cpp.git cd llama.cpp cmake -B build cmake --build build -j --target llama-server llama-cli # Start a local OpenAI-compatible server with a web UI: ./build/bin/llama-server -hf BricksDisplay/jevling-e2b-v1-GGUF:BF16 # Run inference directly in the terminal: ./build/bin/llama-cli -hf BricksDisplay/jevling-e2b-v1-GGUF:BF16
Use Docker
docker model run hf.co/BricksDisplay/jevling-e2b-v1-GGUF:BF16
- LM Studio
- Jan
- Ollama
How to use BricksDisplay/jevling-e2b-v1-GGUF with Ollama:
ollama run hf.co/BricksDisplay/jevling-e2b-v1-GGUF:BF16
- Unsloth Desktop
- Docker Model Runner
How to use BricksDisplay/jevling-e2b-v1-GGUF with Docker Model Runner:
docker model run hf.co/BricksDisplay/jevling-e2b-v1-GGUF:BF16
- Lemonade
How to use BricksDisplay/jevling-e2b-v1-GGUF with Lemonade:
Pull the model
# Download Lemonade from https://lemonade-server.ai/ lemonade pull BricksDisplay/jevling-e2b-v1-GGUF:BF16
Run and chat with the model
lemonade run user.jevling-e2b-v1-GGUF-BF16
List all available models
lemonade list
- Atomic Chat
Jevling-E2B-v1 β GGUF
Quantised build of BricksDisplay/jevling-e2b-v1 for on-device use (16 GB RAM). This is a System One decision model: one forward pass, no generated text, several typed questions per call, each answered as a calibrated probability. It does not work with stock llama.cpp chat/completion endpoints β they can load the weights but cannot ask a typed question or read the answer slot.
Use the maintained implementation
tools/system-one in mybigday/system-one-llama.cpp, branch feat/system-one:
git clone -b feat/system-one https://github.com/mybigday/system-one-llama.cpp
cd system-one-llama.cpp && cmake -B build && cmake --build build --target llama-system-one llama-system-one-batch -j
One call, three typed questions:
build/bin/llama-system-one -m jevling-e2b-v1-q8_0-embf16.gguf \
--state "Customer: I was charged twice for the same subscription this month. Please refund the duplicate charge." \
--choice "queue:Which team should handle this ticket?:billing=payments and refunds,technical=a product fault,account=login or profile settings" \
--noul "refund:Is the customer asking for a refund?" \
--score "urgency:How urgent is this?:routine,soon,urgent,critical" --json
Answers come back as probabilities per option (--json is the same shape the server returns: noul P(true), choice with a distribution, score as the expected level). Pass option descriptions β the model was trained with them. Use --system-one-request request.json for anything with colons or many questions; llama-system-one-batch for throughput. The prompt template and readout live inside the GGUF (tokenizer.chat_template.system_one; the readout is derived from the model, not declared); do not supply a chat template of your own.
Measured: β1.7 s per 5-question request on 16 CPU threads, β2.9 GB RSS; run with --swa-full. Keep flash-attention off on CPU and threads = physical cores.
Or the /v1/systemone endpoint
This file also answers upstream llama.cpp's decision-model endpoint, which needs a build that tokenizes the prompt piece by piece β the boundaries this model was trained on. That is one commit on top of upstream master:
git clone -b system-one/decision-segments https://github.com/mybigday/system-one-llama.cpp
cd system-one-llama.cpp && cmake -B build -DCMAKE_BUILD_TYPE=Release && cmake --build build --target llama-server -j
build/bin/llama-server -m jevling-e2b-v1-q8_0-embf16.gguf -c 2048 -t $(nproc) --n-outputs-max-per-seq 16
curl http://localhost:8080/v1/systemone -H "Content-Type: application/json" -d '{
"state": "Customer: I was charged twice and nobody replied.",
"questions": {
"queue": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "payments and refunds", "technical": "a product fault"}},
"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["routine", "soon", "urgent", "critical"]}
}
}'
Stock llama.cpp must not be used for this. It will load the file and answer, and its answers will be wrong without saying so: a single pass over the whole prompt merges across one of the training boundaries and comes out one token shorter, which moves a probability by up to 0.28 and changes a few decisions in a thousand. Nothing is raised, because the separator the template writes is simply undefined there and renders as empty. The branch above is a requirement, not a suggestion.
Measured against the fp32 reference these weights were validated on, 255 states / 1275 answers: its f32 weights read 1.5e-03 worst-case probability deviation from that reference with 0 of 1275 argmax differences β the figure is larger than the 0.8B's because gemma-4 uses GELU and llama.cpp's CPU GELU goes through an fp16 table, which changes no decision and is upstream's own accepted behaviour, and this file's decision agreement is noul 1.0000 Β· choice 1.0000 Β· score 1.0000, accuracy unchanged on all three, 0 flips.
Evaluation
All numbers are accuracy on datasets the model was not trained on, asked in the System One format (all questions of an item in one request; option descriptions given where the dataset has them). Items per dataset: 120β150 unless noted.
| benchmark | task | Jevling-0.8B-v1 | Jevling-E2B-v1 |
|---|---|---|---|
| MASSIVE (en-US) | scenario classification, 18-way | 0.667 | 0.742 |
| BBC News | topic, 5-way | 0.917 | 0.967 |
| TREC | question type, 6-way | 0.758 | 0.792 |
| PAWS | paraphrase yes/no | 0.717 | 0.700 |
| CommitmentBank | NLI, 3-way | 0.625 | 0.804 |
| StrategyQA | yes/no reasoning | 0.500 | 0.525 |
| PubMedQA | yes/no/maybe | 0.717 | 0.600 |
| SciQ | 4-way science QA | 0.950 | 0.967 |
| Social IQa | 3-way | 0.633 | 0.717 |
| TruthfulQA (MC) | multiple choice | 0.467 | 0.642 |
| XStoryCloze (en) | 2-way | 0.917 | 0.967 |
| QuALITY | long-document 4-way QA | 0.425 | 0.567 |
| RewardBench | pairwise preference | 0.567 | 0.733 |
| Hermes function-calling | tool choice | 0.971 | 0.963 |
| Financial PhraseBank | sentiment, 3-way | 0.658 | 0.667 |
| JevBench easy / original / hard (231 items) | typed decisions | 1.000 / 0.861 / 0.450 | 1.000 / 0.944 / 0.432 |
| zh-TW kiosk set (ours, synthetic-derived, 255 states) | intent acc / completeness AUROC / is-order / noise / size | 0.961 / 0.975 / 0.984 / 0.992 / 1.000 | 0.980 / 0.989 / 0.984 / 0.992 / 1.000 |
JevBench hard (.45) is where every open model we know of sits (TypeSafe's Jev API scores .73 on it); the zh-TW kiosk set is our own synthetic-derived data, so read that row as "fit for the distribution it was built for", not as a general claim.
Limitations
- Long-document, multi-hop, probability and date/time arithmetic questions are weak (JevBench hard .44; QuALITY .4β.6); compute arithmetic in code and put the result in the state.
- Chinese coverage comes from synthetic kiosk-style data; validate on your own transcripts before relying on it.
- When a question is undecidable the model picks the majority-style option; add an explicit "none of the above" option if you need abstention.
- Primarily a decision model, not a chat assistant. The fine-tune only trains the answer slot, so the base model's chat ability is retained (spot-checked in transformers, not benchmarked); this GGUF build is intended for System One use.
Licence
Apache-2.0 (same as the Gemma-4 base). Trained only on commercially usable data: public datasets under MIT, Apache-2.0, CC-BY-4.0, CC-BY-2.0, CC0, ODC-BY and CDLA-Sharing licences, plus our own synthetic data. CC-BY / ODC-BY sources require attribution; the per-dataset list is available on request.
- Downloads last month
- 166
16-bit
Model tree for BricksDisplay/jevling-e2b-v1-GGUF
Base model
google/gemma-4-E2B