Deem-4B: small English language model for local decisions

Deem-4B is one of the best open English decision models at 4B parameters: a small English language model for text classification, routing and decisions with probabilities. Deem-4B: 94.7% accuracy; Kev-4B: 84.7% on held-out English decisions. The test set comes from the same data pipeline as the training data.

Try a decision in the browser, or run the GGUF build locally below.

Run locally

The Q4_K_M build used 3.04 GB RAM at a 4k context. The JevAlt server uses llama.cpp to read decision probabilities and load the matching calibration.

pip install "jevalt[serve,gguf] @ git+https://github.com/mertkayacs/jevalt" && jevalt serve --model mertkayacs/Deem-4B-GGUF --file Deem-4B-Q4_K_M.gguf

Send a situation, question and options to the local server:

import requests

request = {
    "state": "I was charged twice for my subscription. Please refund the duplicate.",
    "questions": {
        "team": {
            "type": "choice",
            "instructions": "Which team should handle this ticket?",
            "criteria": {"billing": "payments, refunds", "technical": "bugs, outages"},
        }
    },
    "reasoning": "off",
    "abstain": False,
}
response = requests.post("http://127.0.0.1:8000/v1/systemone", json=request)
response.raise_for_status()
print(response.json())

Write the situation, question and options in English. The API also supports yes/no questions, ordered scores and an unknown option with abstain: true. The card widget shows a recorded answer from the full-precision model.

Use the full-precision weights with Transformers

This repository holds the full-precision weights. The JevAlt Transformers backend preserves the decision-token format and calibration:

pip install "jevalt[hf] @ git+https://github.com/mertkayacs/jevalt" && jevalt serve --backend hf --model mertkayacs/Deem-4B

See the local run guide for memory and client settings.

Results and limits

On the same held-out English decisions, Deem-4B's probability error (Brier score, lower is better) is 0.091, against 0.256 for Kev-4B. A separate live test on 4 October 2026 scored 122/130 for Deem-4B. Comparison files and live requests, including misses record both tests.

The models were fine-tuned with LoRA from Intern-Decision-4B, using JevAlt training data. English is this model's focus.

  • Refit calibration on your own data before using confidence thresholds. See JevOss.
  • The model accepts text, with an 8k context. Long irrelevant text and date arithmetic remain weak spots.
  • Planted instructions still change some answers. Keep authorization checks outside the model and review consequential decisions.

Downloads and project

GGUF files and measured agreement | Code and training | Deem-4B project page | Training and evaluation details.

Apache-2.0. Cite the JevAlt repository for this release.

An Eschatia Labs project. Mert Kaya.

Downloads last month
177
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Examples
Examples
Billing
0.948
Technical support
0.030
Sales
0.013
Account
0.009
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mertkayacs/Deem-4B

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(3)
this model
Quantizations
2 models

Dataset used to train mertkayacs/Deem-4B

Spaces using mertkayacs/Deem-4B 2

Collections including mertkayacs/Deem-4B