Dinah-0

A 150M-parameter encoder that makes typed decisions, in English and Portuguese, on a CPU.

Dinah-0 reads a question (instructions, state and a set of options) and returns a probability distribution over the options. The same model answers three question types:

Type Output
choice the chosen label, a confidence and the probability of every label
noul the probability that a statement is true
score the expected value on an ordered rubric, plus confidence and probabilities

It does not generate text. It is built for classification, routing, tool selection, extraction checks and similar decisions where a small, fast and cheap model is enough.

Results

Decision Index 0.2.1

Measured with the official kit and scorer (commit 87d4650) and the engine in this repository (decision_index_engine.py at revision e6bfc33), on the full suite (38 benchmarks, 150,759 requests), in bfloat16 on one RTX PRO 4000 Blackwell. Nothing was truncated: the 324 requests longer than 8,192 tokens were refused and count as wrong. Self-reported; the submission to the official leaderboard is pending. Median latency in that run: 5.6 ms per request.

Score
Index (balanced skill) 27.63
Raw index 44.78
Knowledge & Reasoning 0.191
Language Understanding 0.227
Retrieval & Classification 0.478
Tools & Automation 0.379
Arts & Human Taste 0.035

On the live leaderboard of 2026-09-28 this would rank 41st of 72, the best model under 1B parameters (the next one scores 19.22) and the best encoder, ahead of several models between 1.7B and 4.7B parameters.

Other evaluations

Test set n Accuracy
typed-decisions test (English) 2,000 76.35%
typed-decisions test, Portuguese translation 1,895 78.15%
Lichess puzzles, best move among the candidates 2,000 62.00%

Dinah-0 was trained on the typed-decisions train split and on Lichess puzzles (with the test sets held out), so these are in-distribution numbers. The typed-decisions labels come from a teacher model, so accuracy there measures agreement with that teacher.

Speed

Context CPU latency (Apple M5 laptop, int8 ONNX, 4 threads, batch 1)
512 tokens 131 ms
2,048 tokens 0.70 s
8,192 tokens 6.5 s

Measured on an idle machine with an earlier checkpoint of the same architecture (latency depends only on the architecture). Peak memory is ~750 MB at 512 tokens and over 6 GB at 8,192.

Usage

dinah.py (in this repository) has the full input format and two backends with the same API.

pip install torch transformers safetensors huggingface_hub   # PyTorch backend
pip install onnxruntime transformers huggingface_hub         # ONNX int8 backend (CPU)
from huggingface_hub import hf_hub_download
import importlib.util, sys

path = hf_hub_download("Lukitaduarte/dinah-0", "dinah.py")
spec = importlib.util.spec_from_file_location("dinah", path)
dinah = importlib.util.module_from_spec(spec); sys.modules["dinah"] = dinah; spec.loader.exec_module(dinah)

model = dinah.DinahONNX.from_pretrained("Lukitaduarte/dinah-0")   # or dinah.Dinah for PyTorch

model.choice(
    state="The package arrived broken and I need it for tomorrow.",
    instructions="What does the customer want?",
    criteria={
        "refund": "The customer wants their money back",
        "replacement": "The customer wants a new unit sent",
        "complaint": "The customer only wants to complain",
    },
)
# {"type": "choice", "choice": ..., "confidence": ..., "probabilities": {...}}

model.noul(
    state="Hi, please cancel my subscription starting today.",
    instructions="Does the customer want to cancel?",
    true="The customer wants to cancel.",
    false="The customer does not want to cancel.",
)
# {"type": "noul", "noul": ...}
model.score(state="Great product, but shipping took three weeks.", instructions="Rate the review.",
            criteria=["very negative", "negative", "neutral", "positive", "very positive"])

state, instructions and each option can be a string or a JSON object/array. For noul, describe both poles (true and false): with only the statement, the model has a generic "the statement above is false" to compare against and is much less reliable. Several questions can be answered in one call with model.predict([...]).

Input format

Each question becomes one sequence, and each option is scored from the hidden state at its [OPT] marker:

[CLS] instructions [SEP] state [SEP] [OPT] option_1 [OPT] option_2 ... [OPT] option_n

Objects are serialized as compact JSON with sorted keys. For noul, the options are [false, true] and the answer is the probability of true.

Files

File
model.safetensors fp32 weights (ModernBERT encoder + option scorer + confidence head)
config.json ModernBERT config, plus the dinah block ([OPT] token id, question types, max length)
tokenizer.json, tokenizer_config.json, special_tokens_map.json moBERTo tokenizer (original ModernBERT vocabulary) plus the [OPT] token
onnx/model_int8.onnx int8 dynamic quantization with block-local attention, for CPU; needs onnxruntime (uses the com.microsoft MultiHeadAttention op)
dinah.py inference code
decision_index_engine.py engine for the Decision Index kit (--engine decision_index_engine:DinahEngine)

The int8 model agrees with fp32 on the top option in 94.3% of 300 typed-decisions test items, with the same accuracy (73.7% vs 73.0% on that sample). Use the PyTorch weights when you need the exact fp32 probabilities.

Model details

  • Architecture: ModernBERT-base (22 layers, hidden size 768, 8,192-token context), ~150M parameters, with an [OPT] token added to the vocabulary, a question-type embedding, an option scorer (MLP over each [OPT] hidden state) and a small confidence head that reads the shape of the option distribution.
  • Base model: Tropic-AI/moBERTo (Apache-2.0), a ModernBERT further pre-trained on Portuguese, in the variant that keeps the original ModernBERT tokenizer.
  • Training: several fine-tuning stages on typed-decision data (English and Portuguese), long-context multitask data, Lichess puzzles and, in the last stage, 605,256 examples built from the training splits of the Decision Index source datasets (never from the suite's evaluation rows), with Portuguese replay. About 1.19 billion training tokens across all stages.
  • Objective: cross-entropy against the target distribution over options, plus a Brier term; the confidence head is trained to predict whether the top option is correct.

Limitations

  • Decisions only. No generation, no open-ended answers.
  • Little world knowledge. It is weak on questions that need facts it cannot read in the input (see the Knowledge and Arts scores above).
  • Weak at arithmetic over the input. Comparing numbers in the state (totals against limits, dates) is unreliable.
  • Two languages. Trained for English and Portuguese; other languages are untested.
  • 8,192 tokens, no truncation. Longer requests raise an error instead of being silently cut.
  • Confidence is not recalibrated in this release. The confidence value comes straight from the confidence head.
  • Research model. Evaluate it on your own data before relying on it.

License

CC BY-NC 4.0 (non-commercial). Part of the last training stage comes from the MMLU auxiliary training set, which includes RACE, released for non-commercial research only, and MCTest. Because of that, these weights are released for non-commercial use.

Training data and attributions

Datasets in the training path, with their licenses:

Dataset License
ZefanCai/Open-Jev CC0-1.0
LocalLLaMA/typed-decisions and its Portuguese translation (made with Qwen3-4B) Apache-2.0
Lichess puzzles CC0-1.0
ChessBench / searchless_chess CC BY 4.0
MASSIVE CC BY 4.0
Bitext customer support CDLA-Sharing-1.0 (data not redistributed)
ESCI Apache-2.0
WANLI CC BY 4.0
SNLI CC BY-SA 4.0
HoVer CC BY-SA 4.0
CommonsenseQA MIT
HellaSwag MIT (per the original repository)
WinoGrande Apache-2.0
MMLU auxiliary train (ARC, OBQA, RACE, MCTest) MIT (repository); RACE non-commercial research only
GSM8K MIT
CLadder MIT
iSarcasmEval MIT
When2Call CC BY 4.0
glaive-function-calling-v2 Apache-2.0
ToolACE Apache-2.0
HaluEval MIT
IBM Claim Stance CC BY 3.0
CLINC150 CC BY 3.0
BANKING77 CC BY 4.0
ContractNLI CC BY 4.0
CUAD CC BY 4.0
PubMedQA MIT
MedMCQA Apache-2.0
QASC CC BY 4.0
StrategyQA MIT
bAbI-NLI BSD
BIG-bench (tasks outside BBH) Apache-2.0
twitter-financial-news-sentiment MIT
financial-tweets-sentiment MIT
NOSIBLE financial-sentiment ODC-By
FinQA hallucination detection MIT
GoEmotions Apache-2.0
Civil Comments CC0-1.0
B2W-Reviews01 CC BY 4.0
multilingual-sentiments (pt) Apache-2.0
jurisprudencias_br CC BY 4.0
FaQuAD-NLI CC BY 4.0
FAQ BACEN Apache-2.0
ExtraGLUE (RTE and COPA, pt-BR) MIT

Thanks to the authors of all of them, and to the ModernBERT and moBERTo teams for the base model.

Author

Built by Lukita (GitHub) as a personal study project.

Downloads last month
-
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Lukitaduarte/dinah-0

Quantized
(1)
this model