sonda-1.0

GitHub

A 4B decision model for Polish and English. Give it evidence (state), a typed question — yes/no (noul), choice or score — and its options; it returns a calibrated probability for every option from one forward pass, with no generated tokens. It is not a chat or text-generation model.

Use

The model is served by sonda-server (Apache-2.0), which reads the answer from the option letters' logits and applies the per-type temperatures in sonda.conf.

Install the server (Linux with an NVIDIA GPU and CUDA, Python 3.11 or newer, about 10 GB of GPU memory):

git clone https://github.com/sonda-ml/sonda-server.git
cd sonda-server
python -m venv .venv
.venv/bin/pip install -e ".[gpu]"     # the server, PyTorch and transformers

If PyTorch with CUDA is already installed system-wide (an NVIDIA container or an ARM machine), follow the installation notes in the sonda-server README instead.

Download this model and start the server:

huggingface-cli download <this-model-repo> --local-dir models/sonda-1.0
.venv/bin/sonda-server --model models/sonda-1.0 --port 8090

curl -s localhost:8090/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": "Zamówienie 7120 dotarło uszkodzone. Klient prosi o wymianę, nie o zwrot pieniędzy.",
  "questions": {"refund": {"type": "noul", "instructions": "Klient prosi o zwrot pieniędzy."}}}'

The answer has the TypeSafe /v1/systemone shape: noul (P(yes)), choice with probabilities, or score (expected level) with probabilities. Open http://localhost:8090/ for examples to edit and send, and games played by the model.

Or run it with vLLM in Docker — same answers, several times the throughput with many clients. The image builds on the official vLLM image from the sonda-server checkout; the model folder is mounted, not copied:

cd sonda-server
docker build -f docker/Dockerfile.vllm -t sonda-server:vllm .
docker run --rm --device nvidia.com/gpu=all --ipc=host -p 127.0.0.1:8090:8090 \
    -v "$PWD/../models/sonda-1.0:/model:ro" -e SONDA_NAME=sonda-1.0 sonda-server:vllm

--device nvidia.com/gpu=all needs the NVIDIA Container Toolkit (CDI); with the older runtime use --gpus all. vLLM takes SONDA_GPU_MEMORY_UTILIZATION (default 0.25) of the GPU memory for the weights and its cache and needs a few minutes to start (kernel tuning, CUDA graphs). Every server option works as -e SONDA_<OPTION>=...; see the sonda-server README.

Lineage

step weights license
base Qwen3.5-4B (Qwen team) Apache-2.0
JevK5 v0.2 (alibiserikbay/JevK5) Qwen3.5-4B + LoRA, merged Apache-2.0
sonda-0.1 JevK5 v0.2 + LoRA, merged Apache-2.0
sonda-1.0 sonda-0.1 + LoRA, merged Apache-2.0

Training

  • sonda-0.1: 34,901 training records (question set: English, and Polish);
  • sonda-1.0: 1,929 training records (question set: English, and Polish):
  • Calibration: temperature per question type in sonda.json

Evaluation

PriorBench Jev cases (priorbench/jev, MIT; 2,853 distinct requests with a rule-defined answer; 2 October 2026, served by sonda-server). The model was trained on Polish and English only, so French is out of its training languages.

model language records accuracy NLL ECE
jev English 120 98.3% 0.048 0.027
sonda-1.0 English 120 97.5% 0.096 0.025

Our Polish/English bench set: 327 + 327 decision questions (yes/no, choice and score; Polish and English versions of the same cases), the same records for every model. Score = 100 × (0.5 · accuracy + 0.3 · e^−NLL + 0.2 · (1 − ECE)), averaged over language × question-type cells.

model accuracy EN accuracy PL NLL EN / PL ECE EN / PL score
Jev (jev-latest, TypeSafe API, 1 October 2026) 99.1% 97.9% 0.068 / 0.082 0.046 / 0.047 95.9
sonda-pl-1.0 (this model) 97.2% 95.4% 0.101 / 0.138 0.028 / 0.027 93.4
JevK5 v0.3.3 (alibiserikbay/JevK5) 93.3% 92.4% 0.159 / 0.213 0.073 / 0.065 87.7

Jev is TypeSafe's hosted model (version current on the date above); it reports probabilities rounded to two decimals, which slightly affects its NLL and ECE.

Decision Index 0.2.1 (board, kit): 38 public English benchmarks in five areas, about 120,000 requests, each benchmark chance-corrected (0 = random guessing, 100 = perfect). Run on 3 October 2026 with the kit's http engine against sonda-server (vLLM, --typesafe-compat), NVFP4 version of these weights (see below); every request answered.

model Decision Index Knowledge & Reasoning Language Retrieval & Classification Tools & Automation Arts & Taste
Jev 1.13 (TypeSafe) 57.9 51.4 62.0 55.4 75.1 37.7
JPT-4B 43.0 28.7 52.5 45.0 57.2 25.8
sonda-1.0 (NVFP4) 42.6 27.2 48.0 45.7 61.3 28.0
Jet v6.2 42.6 28.7 43.9 48.2 62.9 27.0
Decider 4B 40.7 25.7 46.0 44.7 58.6 25.0
JevK5 v0.3.3 38.8 23.5 43.3 46.2 52.1 27.8

Quantized versions

Made from these weights with AutoRound 0.16 (round-to-nearest; the small in_proj_a/in_proj_b projections of the linear-attention layers stay in bf16). The same tokenizer and sonda.conf: the temperatures fit without recalibration. Measured on our Polish/English bench set (above) on one NVIDIA GB10, 2 October 2026; "same answers" counts the 654 questions answered as by bf16 on the same engine.

Quality:

version engine accuracy EN / PL NLL EN / PL ECE EN / PL score
bf16 sonda-server (transformers) 97.2 / 95.1% 0.101 / 0.140 0.032 / 0.024 93.4
int8 W8A16 sonda-server (transformers) 97.6 / 95.1% 0.101 / 0.141 0.029 / 0.026 93.4
bf16 vLLM 0.29 97.2 / 95.1% 0.100 / 0.141 0.035 / 0.025 93.3
FP8 vLLM 0.29 96.9 / 94.5% 0.095 / 0.130 0.025 / 0.020 93.5
NVFP4 vLLM 0.29 96.6 / 94.5% 0.091 / 0.134 0.025 / 0.034 93.5

Size and speed:

version format weights engine 1 client, median 8 clients
bf16 safetensors 8.4 GB sonda-server (transformers) 304 ms 2.5 req/s
int8 W8A16 AutoRound 4.9 GB sonda-server (transformers) 359 ms 2.6 req/s
bf16 safetensors 8.4 GB vLLM 0.29 275 ms 6.6 req/s
FP8 compressed-tensors 4.8 GB vLLM 0.29 130 ms 11.4 req/s
NVFP4 compressed-tensors 3.3 GB vLLM 0.29 ~86 ms ~18 req/s
  • FP8 and NVFP4 load in vLLM with no extra options (the Docker image above); they need GPU support for the format: FP8 on Ada, Hopper or Blackwell, NVFP4 on Blackwell. They are the fastest versions and lose 0.3–0.6 points of accuracy, while NLL and ECE stay as good as bf16 or better.
  • int8 W8A16 halves the memory with the transformers engine but is not faster (the weights are unpacked for every forward); it needs pip install auto-round.
  • The speed columns are one machine with other jobs on the same GPU; compare them within a column, not with other hardware. The NVFP4 speed was measured on an earlier build with the same kernels.

Limits

  • No arithmetic. The model answers in one pass: it compares values and applies rules well, but does not add, divide, convert units or count days. Compute such values in your code and send the results with the data (a total, an amount per person, a date). Counting words in a list is also weak (about 40–67% on PriorBench).
  • Languages: trained on Polish and English. Other languages work only partly (see French above).
  • Long option lists: one pass reads at most 16 options; runtimes combine several passes for more.

Attribution

Qwen3.5-4B by the Qwen team (Apache-2.0). JevK5 by its authors (Apache-2.0), whose decision prompt and one-pass readout are adapted from SemIf by TheoLeeCJ (MIT). Keep LICENSE and NOTICE with the weights; "Qwen" and "JevK5" name where the model comes from, not this product.

Downloads last month
28
Safetensors
Model size
4B params
Tensor type
BF16
·
F8_E4M3
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Mr-Shmoo/sonda-1.0-4B-FP8

Finetuned
Qwen/Qwen3.5-4B
Quantized
(3)
this model