Noma
Noma is an open decision model from Blackdrome AI Labs. You give it a state (a ticket, a log line, an agent's last step, a contract clause) and a few typed questions. It returns a probability for every option, a separate "none of these" signal, and a measure of its own uncertainty. It never generates text.
- Fast. About 16 ms per decision end to end on one H100.
- Calibrated. Probabilities you can put a threshold on, plus
abstainanduncertainty. - Drop-in. Speaks the same
/v1/systemoneAPI as Jev. - One download. A single
model.safetensors; no base model, no adapter library.
Use it
pip install blackdrome-noma
noma serve # API on http://127.0.0.1:8000/v1/systemone, playground on /
from noma import Noma
model = Noma.from_pretrained("BlackdromeAILabs/noma")
answers, _ = model.decide(
state="Hi, my card was charged twice for order #4471. I also cannot log in since yesterday.",
questions={
"team": {"type": "choice", "instructions": "Which team should handle this first?",
"criteria": {"billing": "Billing and refunds", "identity": "Login and account access",
"shipping": "Shipping and delivery"}},
"refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
},
)
for key, (probs, abstain, uncertainty) in answers.items():
print(key, max(probs, key=probs.get), probs, abstain, uncertainty)
Question types: choice (pick one of your options), noul (yes or no), score (an ordered
scale). Documentation: github.com/blackdromeai-labs/noma.
Results
| Result | |
|---|---|
| Latency, end to end (H100, JevBench client, 231 public tasks) | 16 ms median |
| JevBench easy | 48/48 (100%), ECE 0.011 |
| JevBench original | 71/72 (98.6%), ECE 0.084 |
| Sealed set, 386 human-reviewed decisions, 12 families | 82.6%, ECE 0.042 |
| JevBench, all public tasks | 76.2% |
Cost
$0.023 per 1,000 decisions by JevBench's method (Jev 1.13.0: $0.040). Self-hosted on one H100 at $5.68 an hour, one serial stream: $0.026 per 1,000.
Compared with other decision models
| Model | Base | Easy | Standard | All public | Hard tier | Median latency | Cost per 1k |
|---|---|---|---|---|---|---|---|
| Noma | Qwen3.5-4B, 18 of 32 layers | 100% | 98.6% | 76.2% | 51.4% (46.4% held-out) | 16 ms | $0.023 |
| decider-4b v2 | 4B | 100% | 96.9% | 83.5% | 67.3% | 17 ms | $0.020 |
| Cygnet | frozen Gemma-4-12B | 100% | 96.9% | 87.9% | 75.5% | 35 ms | $0.037 |
| NInfer Flash-Next | large MoE | 100% | 99.0% | 89.6% | 77.3% | 79 ms | |
| JevOne | not disclosed | 100% | 96.9% | 89.6% | 75.0% | 87 ms | |
| Nimble 9B | Qwen3.5-9B | 100% | 94.8% | 79.7% | 65.5% | 389 ms | $0.166 |
| OpenJev (thinking) | 26B MoE, generates reasoning | 100% | 100% | 88.7% | 78.2% | 463 ms | |
| Jev 1.13.0 | not disclosed | 100% | 99.0% | 86.6% | 74.1% | 652 ms | $0.040 |
Other models' figures are from the JevBench leaderboard; Noma's are our own runs with the JevBench client. The hard tier is multi-step reasoning, which Noma leaves to a reasoning model.
Ablations
Every number, the method behind it, per-family breakdowns and the multi-step reasoning tier are in the evaluation document.
Intended use
Noma is the decision layer of a larger system: routing, triage, intent, moderation and
scoring, policy checks, and verifying agent steps (did it work, is the next action safe, is
the task done). It is built to hand off: act automatically when confidence is high, and send
the request to a person or a larger model when abstain or uncertainty is high.
Scope
- Single-pass decisions. Questions that need several chained steps of arithmetic or date reasoning, or tracing a long policy through its amendments, belong with a reasoning model.
- Text only. States up to 4,096 tokens.
- Trained and evaluated in English.
- A decision model informs a decision; for anything with legal, medical, financial or safety consequences, keep a person in the loop.
Model details
| Backbone | Qwen3.5-4B-Base, first 18 of 32 layers, LoRA (rank 16) merged |
| Heads | 4 listwise option scorers (bootstrap ensemble), abstain output, evidence head |
| Input | state + typed questions, with a deterministic fact channel for dates and quantities |
| Output | probabilities per option, abstain probability, ensemble uncertainty |
| Precision | bfloat16, 5.4 GB; runs in 5.2 GB of GPU memory |
| Training data | about 34,000 decisions: permissively licensed public data, code-generated items with exact labels, and items labelled blind by two models with a third as judge |
What Noma introduces: a decision head on a measured mid-depth cut; prefix-fork serving with length buckets and CUDA graphs for a hybrid linear-attention backbone; the fact channel; and an agreement-gated blind labelling cascade. See the architecture document.
Files
| File | |
|---|---|
model.safetensors |
backbone with adapter merged, new-token embeddings, heads |
noma_config.json |
Noma settings |
backbone_config.json |
configuration of the cut backbone |
| tokenizer files | the Qwen tokenizer with Noma's special tokens added |
SHA256SUMS |
checksums |
Citation
@software{noma2026,
title = {Noma: an open, calibrated decision model},
author = {{Blackdrome AI Labs}},
year = {2026},
url = {https://huggingface.co/BlackdromeAILabs/noma}
}
License
MPL-2.0, © Blackdrome AI Labs. Built on Qwen3.5-4B-Base (Apache-2.0). Contact: hello@blackdrome.tech
Model tree for BlackdromeAILabs/noma
Base model
Qwen/Qwen3.5-4B-Base