Instructions to use Mr-Shmoo/sonda-1.0-4B-FP8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Mr-Shmoo/sonda-1.0-4B-FP8 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="Mr-Shmoo/sonda-1.0-4B-FP8") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("Mr-Shmoo/sonda-1.0-4B-FP8") model = AutoModelForCausalLM.from_pretrained("Mr-Shmoo/sonda-1.0-4B-FP8", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use Mr-Shmoo/sonda-1.0-4B-FP8 with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "Mr-Shmoo/sonda-1.0-4B-FP8" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mr-Shmoo/sonda-1.0-4B-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/Mr-Shmoo/sonda-1.0-4B-FP8
- SGLang
How to use Mr-Shmoo/sonda-1.0-4B-FP8 with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "Mr-Shmoo/sonda-1.0-4B-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mr-Shmoo/sonda-1.0-4B-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "Mr-Shmoo/sonda-1.0-4B-FP8" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Mr-Shmoo/sonda-1.0-4B-FP8", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use Mr-Shmoo/sonda-1.0-4B-FP8 with Docker Model Runner:
docker model run hf.co/Mr-Shmoo/sonda-1.0-4B-FP8
sonda-1.0
A 4B decision model for Polish and English. Give it evidence (state), a typed question — yes/no (noul),
choice or score — and its options; it returns a calibrated probability for every option from one forward pass,
with no generated tokens. It is not a chat or text-generation model.
Use
The model is served by sonda-server (Apache-2.0), which reads the
answer from the option letters' logits and applies the per-type temperatures in sonda.conf.
Install the server (Linux with an NVIDIA GPU and CUDA, Python 3.11 or newer, about 10 GB of GPU memory):
git clone https://github.com/sonda-ml/sonda-server.git
cd sonda-server
python -m venv .venv
.venv/bin/pip install -e ".[gpu]" # the server, PyTorch and transformers
If PyTorch with CUDA is already installed system-wide (an NVIDIA container or an ARM machine), follow the installation notes in the sonda-server README instead.
Download this model and start the server:
huggingface-cli download <this-model-repo> --local-dir models/sonda-1.0
.venv/bin/sonda-server --model models/sonda-1.0 --port 8090
curl -s localhost:8090/v1/systemone -H 'Content-Type: application/json' -d '{
"state": "Zamówienie 7120 dotarło uszkodzone. Klient prosi o wymianę, nie o zwrot pieniędzy.",
"questions": {"refund": {"type": "noul", "instructions": "Klient prosi o zwrot pieniędzy."}}}'
The answer has the TypeSafe /v1/systemone shape: noul (P(yes)), choice with probabilities, or score
(expected level) with probabilities. Open http://localhost:8090/ for examples to edit and send, and games played
by the model.
Or run it with vLLM in Docker — same answers, several times the throughput with many clients. The image builds on the official vLLM image from the sonda-server checkout; the model folder is mounted, not copied:
cd sonda-server
docker build -f docker/Dockerfile.vllm -t sonda-server:vllm .
docker run --rm --device nvidia.com/gpu=all --ipc=host -p 127.0.0.1:8090:8090 \
-v "$PWD/../models/sonda-1.0:/model:ro" -e SONDA_NAME=sonda-1.0 sonda-server:vllm
--device nvidia.com/gpu=all needs the NVIDIA Container Toolkit (CDI); with the older runtime use --gpus all.
vLLM takes SONDA_GPU_MEMORY_UTILIZATION (default 0.25) of the GPU memory for the weights and its cache and needs a
few minutes to start (kernel tuning, CUDA graphs). Every server option works as -e SONDA_<OPTION>=...; see the
sonda-server README.
Lineage
| step | weights | license |
|---|---|---|
| base | Qwen3.5-4B (Qwen team) | Apache-2.0 |
JevK5 v0.2 (alibiserikbay/JevK5) |
Qwen3.5-4B + LoRA, merged | Apache-2.0 |
| sonda-0.1 | JevK5 v0.2 + LoRA, merged | Apache-2.0 |
| sonda-1.0 | sonda-0.1 + LoRA, merged | Apache-2.0 |
Training
- sonda-0.1: 34,901 training records (question set: English, and Polish);
- sonda-1.0: 1,929 training records (question set: English, and Polish):
- Calibration: temperature per question type in
sonda.json
Evaluation
PriorBench Jev cases (priorbench/jev, MIT; 2,853 distinct requests with a rule-defined answer; 2 October 2026, served by sonda-server). The model was trained on Polish and English only, so French is out of its training languages.
| model | language | records | accuracy | NLL | ECE |
|---|---|---|---|---|---|
| jev | English | 120 | 98.3% | 0.048 | 0.027 |
| sonda-1.0 | English | 120 | 97.5% | 0.096 | 0.025 |
Our Polish/English bench set: 327 + 327 decision questions (yes/no, choice and score; Polish and English versions of the same cases), the same records for every model. Score = 100 × (0.5 · accuracy + 0.3 · e^−NLL + 0.2 · (1 − ECE)), averaged over language × question-type cells.
| model | accuracy EN | accuracy PL | NLL EN / PL | ECE EN / PL | score |
|---|---|---|---|---|---|
Jev (jev-latest, TypeSafe API, 1 October 2026) |
99.1% | 97.9% | 0.068 / 0.082 | 0.046 / 0.047 | 95.9 |
| sonda-pl-1.0 (this model) | 97.2% | 95.4% | 0.101 / 0.138 | 0.028 / 0.027 | 93.4 |
JevK5 v0.3.3 (alibiserikbay/JevK5) |
93.3% | 92.4% | 0.159 / 0.213 | 0.073 / 0.065 | 87.7 |
Jev is TypeSafe's hosted model (version current on the date above); it reports probabilities rounded to two decimals, which slightly affects its NLL and ECE.
Decision Index 0.2.1 (board,
kit): 38 public English benchmarks in five areas, about 120,000
requests, each benchmark chance-corrected (0 = random guessing, 100 = perfect). Run on 3 October 2026 with the
kit's http engine against sonda-server (vLLM, --typesafe-compat), NVFP4 version of these weights (see below);
every request answered.
| model | Decision Index | Knowledge & Reasoning | Language | Retrieval & Classification | Tools & Automation | Arts & Taste |
|---|---|---|---|---|---|---|
| Jev 1.13 (TypeSafe) | 57.9 | 51.4 | 62.0 | 55.4 | 75.1 | 37.7 |
| JPT-4B | 43.0 | 28.7 | 52.5 | 45.0 | 57.2 | 25.8 |
| sonda-1.0 (NVFP4) | 42.6 | 27.2 | 48.0 | 45.7 | 61.3 | 28.0 |
| Jet v6.2 | 42.6 | 28.7 | 43.9 | 48.2 | 62.9 | 27.0 |
| Decider 4B | 40.7 | 25.7 | 46.0 | 44.7 | 58.6 | 25.0 |
| JevK5 v0.3.3 | 38.8 | 23.5 | 43.3 | 46.2 | 52.1 | 27.8 |
Quantized versions
Made from these weights with AutoRound 0.16 (round-to-nearest; the small
in_proj_a/in_proj_b projections of the linear-attention layers stay in bf16). The same tokenizer and
sonda.conf: the temperatures fit without recalibration. Measured on our Polish/English bench set (above) on one
NVIDIA GB10, 2 October 2026; "same answers" counts the 654 questions answered as by bf16 on the same engine.
Quality:
| version | engine | accuracy EN / PL | NLL EN / PL | ECE EN / PL | score |
|---|---|---|---|---|---|
| bf16 | sonda-server (transformers) | 97.2 / 95.1% | 0.101 / 0.140 | 0.032 / 0.024 | 93.4 |
| int8 W8A16 | sonda-server (transformers) | 97.6 / 95.1% | 0.101 / 0.141 | 0.029 / 0.026 | 93.4 |
| bf16 | vLLM 0.29 | 97.2 / 95.1% | 0.100 / 0.141 | 0.035 / 0.025 | 93.3 |
| FP8 | vLLM 0.29 | 96.9 / 94.5% | 0.095 / 0.130 | 0.025 / 0.020 | 93.5 |
| NVFP4 | vLLM 0.29 | 96.6 / 94.5% | 0.091 / 0.134 | 0.025 / 0.034 | 93.5 |
Size and speed:
| version | format | weights | engine | 1 client, median | 8 clients |
|---|---|---|---|---|---|
| bf16 | safetensors | 8.4 GB | sonda-server (transformers) | 304 ms | 2.5 req/s |
| int8 W8A16 | AutoRound | 4.9 GB | sonda-server (transformers) | 359 ms | 2.6 req/s |
| bf16 | safetensors | 8.4 GB | vLLM 0.29 | 275 ms | 6.6 req/s |
| FP8 | compressed-tensors | 4.8 GB | vLLM 0.29 | 130 ms | 11.4 req/s |
| NVFP4 | compressed-tensors | 3.3 GB | vLLM 0.29 | ~86 ms | ~18 req/s |
- FP8 and NVFP4 load in vLLM with no extra options (the Docker image above); they need GPU support for the format: FP8 on Ada, Hopper or Blackwell, NVFP4 on Blackwell. They are the fastest versions and lose 0.3–0.6 points of accuracy, while NLL and ECE stay as good as bf16 or better.
- int8 W8A16 halves the memory with the transformers engine but is not faster (the weights are unpacked for
every forward); it needs
pip install auto-round. - The speed columns are one machine with other jobs on the same GPU; compare them within a column, not with other hardware. The NVFP4 speed was measured on an earlier build with the same kernels.
Limits
- No arithmetic. The model answers in one pass: it compares values and applies rules well, but does not add, divide, convert units or count days. Compute such values in your code and send the results with the data (a total, an amount per person, a date). Counting words in a list is also weak (about 40–67% on PriorBench).
- Languages: trained on Polish and English. Other languages work only partly (see French above).
- Long option lists: one pass reads at most 16 options; runtimes combine several passes for more.
Attribution
Qwen3.5-4B by the Qwen team (Apache-2.0). JevK5 by its authors (Apache-2.0), whose decision prompt and one-pass readout are adapted from SemIf by TheoLeeCJ (MIT). Keep LICENSE and NOTICE with the weights; "Qwen" and "JevK5" name where the model comes from, not this product.
- Downloads last month
- 28
Model tree for Mr-Shmoo/sonda-1.0-4B-FP8
Base model
Qwen/Qwen3.5-4B-Base