You need to agree to share your contact information to access this model

This repository is publicly accessible, but you have to accept the conditions to access its files and content.

Log in or Sign Up to review the conditions and access this model content.

Snap 3B (BF16)

Snap is a small, fast System One decision model, in the style of TypeSafe's Jev. You give it a piece of state and typed questions (noul yes/no, choice, score). For every allowed answer it returns a calibrated probability, read from a single forward pass. It never generates text, so its answers can't be malformed.

  • Base model: mistralai/Ministral-3-3B-Instruct-2512-BF16 (Mistral AI), LoRA fine-tuned and merged into the base weights.
  • Training data: weighted towards banking and financial services (Banking77, CFPB complaint narratives, UCI credit and bank-marketing data), plus general decision tasks and synthetic policy and routing scenarios.
  • Serving: stock vLLM, with no patches or custom kernels. Also available as FP8, which has the same accuracy and is about 30% faster.
  • JevBench: 0.727 on the 231 public items, against Jev 1.13.0's 0.866 and 0.710–0.714 for Kev 8B and Decider 2B.

Snap is an independent project. It is not affiliated with or endorsed by TypeSafe AI or Mistral AI.

How to use

Snap is a normal causal LM checkpoint, but it's only meaningful through its label-token readout: every question is rendered in a fixed prompt format, and the probability of each answer is read from the single next token. Use the gateway from the Snap code repository (jevlite/), which exposes the Jev-compatible POST /v1/systemone API:

vllm serve bm-natwest/snap-3b --served-model-name jevlite --max-model-len 16384 \
  --enable-prefix-caching --logprobs-mode processed_logprobs --max-logprobs 32 \
  --limit-mm-per-prompt '{"image":0}'
JEVLITE_MODEL=bm-natwest/snap-3b JEVLITE_CONFIG=jevlite_config.json \
  uvicorn jevlite.gateway:app --port 8080
POST /v1/systemone
{"state": "I topped up by bank transfer on Monday and it still is not in my balance.",
 "model": "snap",
 "questions": {
   "intent": {"type": "choice", "instructions": "Which banking intent?",
              "criteria": {"pending_top_up": "Top-up not showing yet", "card_arrival": "Waiting for a card", "other": "Something else"}},
   "urgent": {"type": "noul", "instructions": "Is the request time-sensitive?"}}}

Prompt format (defined in jevlite/prompt.py):

  • A fixed system prompt, then a user turn: <state>…</state>, <question type=…>…</question>, <options>…</options>.
  • The answer is one label token: A–Z for choice, 0–9 for score, Y/N for noul.
  • Serving reads the next-token logprobs restricted to those labels (allowed_token_ids) at the fitted temperature.
  • Temperatures (jevlite_config.json, fitted on validation): choice 0.9412, noul 0.9192, score 1.023.
  • Choices with more than 26 options are scored in chunks of 26, then a final selection runs over the finalists.
  • SYSTEM_PROMPT.txt in this repo is the base model's file, carried over unchanged. Snap uses its own system prompt, the one in jevlite/prompt.py.

Results

JevBench public items

The 231 public decisions of JevBench (harness pinned in third_party/jevbench.COMMIT), scored by the harness's own code, one request at a time. Reference rows are the published per-item outcomes (results/v1.2/jevbench-v1.2-per-task.json). The official board also scores held-out and sealed items plus speed and cost, so these are not official board scores.

System Easy (48) Standard (72) Hard (111) All (231) ECE p50 ms
Jev 1.13.0 (TypeSafe AI) 1.000 0.986 0.730 0.866
djev (Maisa, diffusion-gemma) 1.000 0.986 0.676 0.840
Snap v1 BF16 1.000 0.889 0.505 0.727 0.130 50
Snap v1 FP8 1.000 0.903 0.495 0.727 0.135 34
kev 8B (research preview) 1.000 0.931 0.450 0.714
decider-2b (Mapika) 1.000 0.847 0.495 0.710
Ministral 3 3B, untrained 0.979 0.806 0.468 0.680 0.101 48
kev 4B (research preview) 1.000 0.889 0.369 0.662
smalljev semantic-v9 0.979 0.681 0.396 0.606
Laya (Convai Innovations, ModernBERT-large 421M) 0.958 0.694 0.351 0.584
openJev Verdict (heman10x, ModernBERT-base 151M) 0.854 0.625 0.378 0.554
Laya-style 400M, warm-up + finance 0.979 0.639 0.315 0.554 0.160 17
Laya-style 400M, finance only 0.896 0.625 0.342 0.545 0.118 54
kev 0.5B 0.958 0.486 0.297 0.494

Our evaluation sets (macro over families)

Model Validation (mixed) Public test (18 families) Held-out teacher domains Programmatic test
Ministral 3 3B, untrained (T=1, NGC 26.04 stack) 0.694 / 0.205 0.506 / 0.186 0.778 / 0.063 0.752 / 0.165
Snap v1 BF16, T=1 0.853 / 0.103 0.740 / 0.054 0.868 / 0.056 0.935 / 0.065
Snap v1 BF16, calibrated 0.855 / 0.100 0.740 / 0.054 0.869 / 0.053 0.934 / 0.061
Snap v1 FP8, calibrated 0.852 / 0.104 0.738 / 0.053 0.866 / 0.065 0.934 / 0.062

Temperatures fitted on validation for v1: see runs/v1/jevlite_config.json.

Finance families (public test split)

Family Untrained 3B Snap v1 BF16 Snap v1 FP8 Laya-style 400M finance-only Laya-style 400M warm-up+finance
banking77 0.607 / 0.117 0.829 / 0.053 0.828 / 0.050 0.925 / 0.047 0.866 / 0.076
cfpb_product 0.668 / 0.175 0.839 / 0.038 0.840 / 0.032 0.844 / 0.037 0.812 / 0.037
cfpb_issue 0.414 / 0.278 0.619 / 0.051 0.611 / 0.047 0.673 / 0.061 0.640 / 0.048
german_credit 0.349 / 0.427 0.709 / 0.095 0.721 / 0.070 0.709 / 0.058 0.698 / 0.070
bank_marketing 0.882 / 0.047 0.897 / 0.032 0.896 / 0.027 0.891 / 0.021 0.898 / 0.027
macro 0.584 / 0.209 0.779 / 0.054 0.779 / 0.045 0.808 / 0.045 0.783 / 0.051

Training

  • LoRA: r=32 and α=64 on every linear layer of the language model; the vision tower is frozen. 2 epochs, 46,538 examples.
  • Loss: KL to the soft target plus 0.5 × Brier, computed over the question's label tokens only. One-hot targets are smoothed by 0.05.
  • Calibration: one post-hoc temperature per question type.
Source Examples Notes
Public datasets, from the original publishers 25,643 Banking77 (PolyAI, CC-BY-4.0); CFPB complaint narratives (US public domain); UCI Statlog German Credit and UCI Bank Marketing (CC-BY-4.0); CLINC150 (CC-BY-3.0); LexGLUE LEDGAR (CC-BY-4.0); GoEmotions (Apache-2.0); Measuring Hate Speech (CC-BY-4.0); Civil Comments (CC0); HelpSteer2 (CC-BY-4.0); BoolQ (CC-BY-SA-3.0); PAWS; SNLI (CC-BY-SA-4.0); MMLU auxiliary train (MIT); ARC-Challenge (CC-BY-SA-4.0)
Synthetic scenarios by gpt-oss-120b (OpenAI) 13,186 40 domains, 10 state formats, 7 EU languages; labelled from 2–4 teacher votes with shuffled option order
Programmatic rules 7,709 Exact answers: policies, thresholds, dates, counts, long-context needles, prompt injection

Limitations

  • Hard reasoning: multi-step reasoning over long states (JevBench's hard tier, 0.50 against Jev's 0.73) is the main weakness. General knowledge is limited by the 3B base.
  • Calibration on unfamiliar items: it's well calibrated on data like its training data (ECE about 0.05), but over-confident on unfamiliar hard items (ECE 0.13 on JevBench).
  • Not for automated decisions without human oversight where errors carry legal, financial or safety consequences. Use the probabilities to route uncertain cases to people.
  • NVFP4: an NVFP4 build is not provided; it gave incorrect outputs on vLLM 0.27 on the GB10.

Licence and attribution

  • Licence: not yet declared for this fine-tune; to be decided by the owner.
  • Base model: Ministral 3 3B is © Mistral AI under Apache-2.0, and its licence and notices apply to the base-model portion of these weights.
  • Training data: each dataset listed above keeps its own licence and attribution.
Downloads last month
3
Safetensors
Model size
4B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bm-natwest/snap-3b