Noma: decisions in milliseconds. An open decision model by Blackdrome AI Labs.

Noma

Noma is an open decision model from Blackdrome AI Labs. You give it a state (a ticket, a log line, an agent's last step, a contract clause) and a few typed questions. It returns a probability for every option, a separate "none of these" signal, and a measure of its own uncertainty. It never generates text.

  • Fast. About 16 ms per decision end to end on one H100.
  • Calibrated. Probabilities you can put a threshold on, plus abstain and uncertainty.
  • Drop-in. Speaks the same /v1/systemone API as Jev.
  • One download. A single model.safetensors; no base model, no adapter library.

The Noma playground

Use it

pip install blackdrome-noma
noma serve          # API on http://127.0.0.1:8000/v1/systemone, playground on /
from noma import Noma

model = Noma.from_pretrained("BlackdromeAILabs/noma")
answers, _ = model.decide(
    state="Hi, my card was charged twice for order #4471. I also cannot log in since yesterday.",
    questions={
        "team": {"type": "choice", "instructions": "Which team should handle this first?",
                 "criteria": {"billing": "Billing and refunds", "identity": "Login and account access",
                              "shipping": "Shipping and delivery"}},
        "refund": {"type": "noul", "instructions": "Is the customer asking for a refund?"},
    },
)
for key, (probs, abstain, uncertainty) in answers.items():
    print(key, max(probs, key=probs.get), probs, abstain, uncertainty)

Question types: choice (pick one of your options), noul (yes or no), score (an ordered scale). Documentation: github.com/blackdromeai-labs/noma.

Results

Median latency per decision

Result
Latency, end to end (H100, JevBench client, 231 public tasks) 16 ms median
JevBench easy 48/48 (100%), ECE 0.011
JevBench original 71/72 (98.6%), ECE 0.084
Sealed set, 386 human-reviewed decisions, 12 families 82.6%, ECE 0.042
JevBench, all public tasks 76.2%

Accuracy

Cost

Cost per 1,000 decisions

$0.023 per 1,000 decisions by JevBench's method (Jev 1.13.0: $0.040). Self-hosted on one H100 at $5.68 an hour, one serial stream: $0.026 per 1,000.

Compared with other decision models

Model Base Easy Standard All public Hard tier Median latency Cost per 1k
Noma Qwen3.5-4B, 18 of 32 layers 100% 98.6% 76.2% 51.4% (46.4% held-out) 16 ms $0.023
decider-4b v2 4B 100% 96.9% 83.5% 67.3% 17 ms $0.020
Cygnet frozen Gemma-4-12B 100% 96.9% 87.9% 75.5% 35 ms $0.037
NInfer Flash-Next large MoE 100% 99.0% 89.6% 77.3% 79 ms
JevOne not disclosed 100% 96.9% 89.6% 75.0% 87 ms
Nimble 9B Qwen3.5-9B 100% 94.8% 79.7% 65.5% 389 ms $0.166
OpenJev (thinking) 26B MoE, generates reasoning 100% 100% 88.7% 78.2% 463 ms
Jev 1.13.0 not disclosed 100% 99.0% 86.6% 74.1% 652 ms $0.040

Other models' figures are from the JevBench leaderboard; Noma's are our own runs with the JevBench client. The hard tier is multi-step reasoning, which Noma leaves to a reasoning model.

Ablations

Ablations

Every number, the method behind it, per-family breakdowns and the multi-step reasoning tier are in the evaluation document.

Intended use

Noma is the decision layer of a larger system: routing, triage, intent, moderation and scoring, policy checks, and verifying agent steps (did it work, is the next action safe, is the task done). It is built to hand off: act automatically when confidence is high, and send the request to a person or a larger model when abstain or uncertainty is high.

Scope

  • Single-pass decisions. Questions that need several chained steps of arithmetic or date reasoning, or tracing a long policy through its amendments, belong with a reasoning model.
  • Text only. States up to 4,096 tokens.
  • Trained and evaluated in English.
  • A decision model informs a decision; for anything with legal, medical, financial or safety consequences, keep a person in the loop.

Model details

Noma architecture

Backbone Qwen3.5-4B-Base, first 18 of 32 layers, LoRA (rank 16) merged
Heads 4 listwise option scorers (bootstrap ensemble), abstain output, evidence head
Input state + typed questions, with a deterministic fact channel for dates and quantities
Output probabilities per option, abstain probability, ensemble uncertainty
Precision bfloat16, 5.4 GB; runs in 5.2 GB of GPU memory
Training data about 34,000 decisions: permissively licensed public data, code-generated items with exact labels, and items labelled blind by two models with a third as judge

Decision head

What Noma introduces: a decision head on a measured mid-depth cut; prefix-fork serving with length buckets and CUDA graphs for a hybrid linear-attention backbone; the fact channel; and an agreement-gated blind labelling cascade. See the architecture document.

Files

File
model.safetensors backbone with adapter merged, new-token embeddings, heads
noma_config.json Noma settings
backbone_config.json configuration of the cut backbone
tokenizer files the Qwen tokenizer with Noma's special tokens added
SHA256SUMS checksums

Citation

@software{noma2026,
  title  = {Noma: an open, calibrated decision model},
  author = {{Blackdrome AI Labs}},
  year   = {2026},
  url    = {https://huggingface.co/BlackdromeAILabs/noma}
}

License

MPL-2.0, © Blackdrome AI Labs. Built on Qwen3.5-4B-Base (Apache-2.0). Contact: hello@blackdrome.tech

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
3B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for BlackdromeAILabs/noma

Finetuned
(208)
this model