Instructions to use bm-natwest/snap-3b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use bm-natwest/snap-3b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="bm-natwest/snap-3b")# pip install -U transformers accelerate # Load model directly from transformers import AutoProcessor, AutoModelForMultimodalLM processor = AutoProcessor.from_pretrained("bm-natwest/snap-3b") model = AutoModelForMultimodalLM.from_pretrained("bm-natwest/snap-3b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Snap 3B (BF16)
Snap is a small, fast System One decision model, in the style of TypeSafe's Jev. You give
it a piece of state and typed questions (noul yes/no, choice, score). For every allowed
answer it returns a calibrated probability, read from a single forward pass. It never
generates text, so its answers can't be malformed.
- Base model: mistralai/Ministral-3-3B-Instruct-2512-BF16 (Mistral AI), LoRA fine-tuned and merged into the base weights.
- Training data: weighted towards banking and financial services (Banking77, CFPB complaint narratives, UCI credit and bank-marketing data), plus general decision tasks and synthetic policy and routing scenarios.
- Serving: stock vLLM, with no patches or custom kernels. Also available as FP8, which has the same accuracy and is about 30% faster.
- JevBench: 0.727 on the 231 public items, against Jev 1.13.0's 0.866 and 0.710–0.714 for Kev 8B and Decider 2B.
Snap is an independent project. It is not affiliated with or endorsed by TypeSafe AI or Mistral AI.
How to use
Snap is a normal causal LM checkpoint, but it's only meaningful through its label-token
readout: every question is rendered in a fixed prompt format, and the probability of each
answer is read from the single next token. Use the gateway from the Snap code repository
(jevlite/), which exposes the Jev-compatible POST /v1/systemone API:
vllm serve bm-natwest/snap-3b --served-model-name jevlite --max-model-len 16384 \
--enable-prefix-caching --logprobs-mode processed_logprobs --max-logprobs 32 \
--limit-mm-per-prompt '{"image":0}'
JEVLITE_MODEL=bm-natwest/snap-3b JEVLITE_CONFIG=jevlite_config.json \
uvicorn jevlite.gateway:app --port 8080
POST /v1/systemone
{"state": "I topped up by bank transfer on Monday and it still is not in my balance.",
"model": "snap",
"questions": {
"intent": {"type": "choice", "instructions": "Which banking intent?",
"criteria": {"pending_top_up": "Top-up not showing yet", "card_arrival": "Waiting for a card", "other": "Something else"}},
"urgent": {"type": "noul", "instructions": "Is the request time-sensitive?"}}}
Prompt format (defined in jevlite/prompt.py):
- A fixed system prompt, then a user turn:
<state>…</state>,<question type=…>…</question>,<options>…</options>. - The answer is one label token:
A–Zforchoice,0–9forscore,Y/Nfornoul. - Serving reads the next-token logprobs restricted to those labels (
allowed_token_ids) at the fitted temperature. - Temperatures (
jevlite_config.json, fitted on validation): choice 0.9412, noul 0.9192, score 1.023. - Choices with more than 26 options are scored in chunks of 26, then a final selection runs over the finalists.
SYSTEM_PROMPT.txtin this repo is the base model's file, carried over unchanged. Snap uses its own system prompt, the one injevlite/prompt.py.
Results
JevBench public items
The 231 public decisions of JevBench (harness pinned in third_party/jevbench.COMMIT), scored by the harness's own code, one request at a time. Reference rows are the published per-item outcomes (results/v1.2/jevbench-v1.2-per-task.json). The official board also scores held-out and sealed items plus speed and cost, so these are not official board scores.
| System | Easy (48) | Standard (72) | Hard (111) | All (231) | ECE | p50 ms |
|---|---|---|---|---|---|---|
| Jev 1.13.0 (TypeSafe AI) | 1.000 | 0.986 | 0.730 | 0.866 | ||
| djev (Maisa, diffusion-gemma) | 1.000 | 0.986 | 0.676 | 0.840 | ||
| Snap v1 BF16 | 1.000 | 0.889 | 0.505 | 0.727 | 0.130 | 50 |
| Snap v1 FP8 | 1.000 | 0.903 | 0.495 | 0.727 | 0.135 | 34 |
| kev 8B (research preview) | 1.000 | 0.931 | 0.450 | 0.714 | ||
| decider-2b (Mapika) | 1.000 | 0.847 | 0.495 | 0.710 | ||
| Ministral 3 3B, untrained | 0.979 | 0.806 | 0.468 | 0.680 | 0.101 | 48 |
| kev 4B (research preview) | 1.000 | 0.889 | 0.369 | 0.662 | ||
| smalljev semantic-v9 | 0.979 | 0.681 | 0.396 | 0.606 | ||
| Laya (Convai Innovations, ModernBERT-large 421M) | 0.958 | 0.694 | 0.351 | 0.584 | ||
| openJev Verdict (heman10x, ModernBERT-base 151M) | 0.854 | 0.625 | 0.378 | 0.554 | ||
| Laya-style 400M, warm-up + finance | 0.979 | 0.639 | 0.315 | 0.554 | 0.160 | 17 |
| Laya-style 400M, finance only | 0.896 | 0.625 | 0.342 | 0.545 | 0.118 | 54 |
| kev 0.5B | 0.958 | 0.486 | 0.297 | 0.494 |
Our evaluation sets (macro over families)
| Model | Validation (mixed) | Public test (18 families) | Held-out teacher domains | Programmatic test |
|---|---|---|---|---|
| Ministral 3 3B, untrained (T=1, NGC 26.04 stack) | 0.694 / 0.205 | 0.506 / 0.186 | 0.778 / 0.063 | 0.752 / 0.165 |
| Snap v1 BF16, T=1 | 0.853 / 0.103 | 0.740 / 0.054 | 0.868 / 0.056 | 0.935 / 0.065 |
| Snap v1 BF16, calibrated | 0.855 / 0.100 | 0.740 / 0.054 | 0.869 / 0.053 | 0.934 / 0.061 |
| Snap v1 FP8, calibrated | 0.852 / 0.104 | 0.738 / 0.053 | 0.866 / 0.065 | 0.934 / 0.062 |
Temperatures fitted on validation for v1: see runs/v1/jevlite_config.json.
Finance families (public test split)
| Family | Untrained 3B | Snap v1 BF16 | Snap v1 FP8 | Laya-style 400M finance-only | Laya-style 400M warm-up+finance |
|---|---|---|---|---|---|
| banking77 | 0.607 / 0.117 | 0.829 / 0.053 | 0.828 / 0.050 | 0.925 / 0.047 | 0.866 / 0.076 |
| cfpb_product | 0.668 / 0.175 | 0.839 / 0.038 | 0.840 / 0.032 | 0.844 / 0.037 | 0.812 / 0.037 |
| cfpb_issue | 0.414 / 0.278 | 0.619 / 0.051 | 0.611 / 0.047 | 0.673 / 0.061 | 0.640 / 0.048 |
| german_credit | 0.349 / 0.427 | 0.709 / 0.095 | 0.721 / 0.070 | 0.709 / 0.058 | 0.698 / 0.070 |
| bank_marketing | 0.882 / 0.047 | 0.897 / 0.032 | 0.896 / 0.027 | 0.891 / 0.021 | 0.898 / 0.027 |
| macro | 0.584 / 0.209 | 0.779 / 0.054 | 0.779 / 0.045 | 0.808 / 0.045 | 0.783 / 0.051 |
Training
- LoRA: r=32 and α=64 on every linear layer of the language model; the vision tower is frozen. 2 epochs, 46,538 examples.
- Loss: KL to the soft target plus 0.5 × Brier, computed over the question's label tokens only. One-hot targets are smoothed by 0.05.
- Calibration: one post-hoc temperature per question type.
| Source | Examples | Notes |
|---|---|---|
| Public datasets, from the original publishers | 25,643 | Banking77 (PolyAI, CC-BY-4.0); CFPB complaint narratives (US public domain); UCI Statlog German Credit and UCI Bank Marketing (CC-BY-4.0); CLINC150 (CC-BY-3.0); LexGLUE LEDGAR (CC-BY-4.0); GoEmotions (Apache-2.0); Measuring Hate Speech (CC-BY-4.0); Civil Comments (CC0); HelpSteer2 (CC-BY-4.0); BoolQ (CC-BY-SA-3.0); PAWS; SNLI (CC-BY-SA-4.0); MMLU auxiliary train (MIT); ARC-Challenge (CC-BY-SA-4.0) |
| Synthetic scenarios by gpt-oss-120b (OpenAI) | 13,186 | 40 domains, 10 state formats, 7 EU languages; labelled from 2–4 teacher votes with shuffled option order |
| Programmatic rules | 7,709 | Exact answers: policies, thresholds, dates, counts, long-context needles, prompt injection |
Limitations
- Hard reasoning: multi-step reasoning over long states (JevBench's hard tier, 0.50 against Jev's 0.73) is the main weakness. General knowledge is limited by the 3B base.
- Calibration on unfamiliar items: it's well calibrated on data like its training data (ECE about 0.05), but over-confident on unfamiliar hard items (ECE 0.13 on JevBench).
- Not for automated decisions without human oversight where errors carry legal, financial or safety consequences. Use the probabilities to route uncertain cases to people.
- NVFP4: an NVFP4 build is not provided; it gave incorrect outputs on vLLM 0.27 on the GB10.
Licence and attribution
- Licence: not yet declared for this fine-tune; to be decided by the owner.
- Base model: Ministral 3 3B is © Mistral AI under Apache-2.0, and its licence and notices apply to the base-model portion of these weights.
- Training data: each dataset listed above keeps its own licence and attribution.
- Downloads last month
- 3
Model tree for bm-natwest/snap-3b
Base model
mistralai/Ministral-3-3B-Base-2512