BaseDecision Model Card
BaseDecision is a compact, text-only decision model for selecting among caller-provided options, answering boolean questions, and assigning discrete ratings. It combines a ModernBERT-large encoder and additional task adaptation. It processes up to 8,192 packed tokens per request and is approximately 0.42 billion parameters.
BaseDecision received approximately 320 million tokens of additional training. In the supplied third-party evaluation, it achieved the highest unweighted average score among the five compared models and led five of eight benchmarks.
Model details
| Property | Description |
|---|---|
| Model name | BaseDecision |
| Developer | Hrudayaditya βAadyβ Jallu |
| Repository | hrudayaditya/BaseDecision |
| Model family | ModernBERT-large encoder |
| Approximate size | 0.42B parameters; rounded model size, not an exact parameter inventory |
| Architecture | Bidirectional transformer encoder with typed decision heads |
| Input | Text context, question/instructions, and candidate options or rating levels |
| Output | Selected answer, option ID, raw logits, option probabilities, and request metadata |
| Decision types | choice, noul (boolean), and score (discrete rating) |
| Candidate options | 2β255, subject to the total token budget |
| Context limit | 8,192 packed tokens, including context, instructions, options, and special tokens |
| Primary evaluated language | English |
| Additional training | Approximately 320M tokens |
| Public Python package | basedecision; published version reviewed: 0.1.1, Python 3.10+ |
This is a discriminative decision model. It does not generate free-form explanations, expose a causal language-model .generate() interface, or provide the general-purpose capabilities of a chat model. Multiple questions over one context are evaluated independently; shared transformer-state computation is not claimed.
Intended uses
- Intent classification and workflow routing with explicit candidate actions.
- Entity-specific sentiment and stance classification.
- Textual entailment and document-based decision support.
- Boolean checks and discrete ratings over supplied text.
- Local applications that need structured decisions without a remote API call.
Plain-text option labels are supported. Descriptions can help disambiguate labels but are not mandatory. Evaluate the actual question wording, options, input lengths, and data distribution used by the application.
The model should not be treated as an autonomous authority for medical, legal, financial, or security decisions. A correct output schema does not establish factual correctness, safe action selection, or suitability for a high-impact workflow.
Training
Training volume
Approximately 320 million tokens were processed during BaseDecision training, as verified by the developer. This is the additional BaseDecision training volume, not the upstream ModernBERT pretraining corpus size.
Third-party evaluation
The table below reproduces the third-party evaluation results supplied by the developer and published in the repository's benchmark chart. Higher is better. Values are reported benchmark scores; the chart does not identify every row's metric, so they are not uniformly labeled accuracy or F1 here.
| Benchmark | GLiNER 2.5 base | GLiNER2.5-Decide | Decision 1.0 Kai 0.6B | Laya | BaseDecision |
|---|---|---|---|---|---|
| BANKING77 | 23.8 | 65.6 | 40.7 | 14.3 | 68.2 |
| FinEntity | 70.1 | 66.2 | 37.0 | 61.0 | 71.3 |
| ContractNLI | 22.9 | 21.8 | 36.6 | 29.0 | 66.8 |
| VAST | 35.8 | 35.3 | 20.8 | 40.5 | 66.2 |
| SGD / SGD-X | 45.9 | 0.8 | 48.5 | 42.4 | 51.1 |
| MuSR | 34.8 | 45.2 | 45.2 | 43.2 | 44.6 |
| NLI4CT | 39.8 | 48.0 | 29.8 | 47.7 | 42.3 |
| PhishNChips | 50.4 | 50.0 | 49.9 | 50.1 | 50.0 |
| Unweighted average | 40.4 | 41.6 | 38.6 | 41.0 | 57.6 |
BaseDecision leads 5 of 8 benchmarks in this comparison: BANKING77, FinEntity, ContractNLI, VAST, and SGD/SGD-X. Its reported average is 16.0 points above the next-highest reported average,
Confidence and calibration
Local inference returns raw softmax probabilities by default. These are relative scores over the supplied choices, not guaranteed probabilities that an answer is correct. Changing the candidate set can change those scores. A high score does not establish evidence completeness or eliminate the possibility of a confidently wrong answer.
Optional scalar-temperature profiles are available for specific assessed workloads:
| Profile | Release behavior | Temperature | Assessment NLL, raw β scaled | Equal-width ECE, raw β scaled |
|---|---|---|---|---|
| SGD identifier | Explicit opt-in | 1.4646 | 0.6116 β 0.5830 | 8.05% β 2.97% |
| SGD schema | Explicit opt-in | 1.4793 | 0.6034 β 0.5606 | 7.88% β 2.96% |
| VAST | Raw recommended; negligible gain | 1.0071 | 0.6246 β 0.6242 | 7.68% β 7.43% |
These are internal calibration assessments, separate from the third-party benchmark table. Each SGD format was assessed on 1,270 questions from 125 dialogue groups; the formats share groups. VAST used 118 questions/32 groups; CLINC used 156 questions/groups. SGD's short and two-option slices regressed despite aggregate gains.
Deployment and usage
Install the published Python package:
pip install basedecision
Version 0.1.1 is published on PyPI. Python 3.10 or newer is required. The default installation includes PyTorch, Transformers, SafeTensors, Tokenizers, Hugging Face Hub, OpenAI, and Anthropic dependencies. The initial download can therefore be large. To reproduce the reviewed package release, use pip install basedecision==0.1.1; pin the runtime and model revision separately when exact reproducibility matters.
The older runtime, hub, and provider extras remain compatibility aliases; they are no longer needed to install those dependencies. Advanced users who already manage their runtime can use pip install --no-deps basedecision and install the dependencies they need themselves. From a repository clone, use pip install ..
Installing the package does not install model weights. Supply a local checkpoint folder. The README's Hugging Face model identifier remains a placeholder; load_from_hub(repository_id, revision=commit_hash) supports downloading a separately published checkpoint, with local_files_only=True available for cached operation. Do not use the placeholder as an actual model ID.
Load an exported checkpoint containing model.safetensors, its encoder configuration, decision configuration, and tokenizer:
from basedecision import load
model = load('/path/to/model')
result = model.choose(
context='I was charged twice. Please refund the duplicate payment.',
question='What does the customer request?',
options=['Refund request', 'Delivery status', 'Change address'],
)
print(result.answer)
print(result.probabilities) # Raw softmax by default; not calibrated confidence.
model.close()
The current public SDK selects a suitable CUDA path when available and otherwise uses CPU FP32. Explicit device choices are supported. Apple-silicon Macs use the CPU path; native MPS acceleration is not claimed. The opt-in cpu_fast backend uses BF16 and local-attention optimizations, is experimental, checks compatibility at load time, and may be substantially slower on unsuitable hardware. It does not silently substitute another backend. The normal runtime allows Transformers >=4.48,<6, while cpu_fast supports only its checked Transformers 4.48β4.57 range. A default installation can therefore resolve to Transformers 5 and work with the normal backend while cpu_fast raises CPUFastUnavailable. Use a compatible runtime deliberately if testing this experimental path.
Inputs are packed with intact options; over-budget requests raise ContextLengthError rather than being silently truncated. The 8,192-token limit covers the entire packed request, not 8,192 context tokens plus unlimited instructions and options. Simple chunking or voting does not guarantee preservation of cross-document reasoning.
OpenAI and Anthropic adapters provide an alternative SDK backend. They call those providers' models, not BaseDecision weights; their results must not be attributed to BaseDecision. They require explicit selection and send inputs to the chosen provider. Cloud results do not supply BaseDecision's local option probabilities.
Python interfaces and resource management
| Interface | Purpose |
|---|---|
model.choose(...) |
Select an option from plain strings or Option objects |
model.check(...) |
Return a boolean decision |
model.score(...) |
Select a rating level; optionally return its numeric value and expected value |
decide(model, context=..., questions=...) |
Answer several typed questions independently |
model.predict_batch(...) / model.predict_iter(...) |
Process request collections or streams |
result.to_dict() |
Export a result as ordinary JSON-compatible values |
model.count_tokens(request) |
Inspect the packed request length before inference |
Use with load('/path/to/model') as model: or call model.close() to release resources. Closing is idempotent; subsequent prediction on a closed model raises InputError.
The SystemOne adapter accepts the Jev/SystemOne state and typed questions wire format. It is a local-model interface with one forward pass per question; it does not make BaseDecision a Jev model or add shared-context computation. A reference HTTP server is provided in the repository's examples/ directory and binds to localhost. Review SystemOne documentation before exposing it over a network.
Errors and privacy
ContextLengthErrorreports an over-budget input, including required and maximum token counts.InputErrorreports invalid request types, duplicate options, or other malformed input. Context must be text; serialize structured data explicitly.CPUFastUnavailablereports an unsupported or failed experimental CPU configuration.- Cloud failures raise
ProviderErrorwithcode,retryable, andstatus_code. The documented error representation excludes API keys and request text. - Refused or invalid provider responses raise errors rather than producing a guessed label or silently returning
False.
Local inference requires no provider API key. Explicit cloud-backend selection sends text to that provider; there is no automatic cloud fallback. See the API reference, cloud documentation, and security guidance.
Performance and hardware
Latency depends on sequence length, option count, batching, processor instruction support, runtime, and system load. No universal CPU latency guarantee is made.
The repository reports these warmed, synthetic, similar-length throughput measurements on an A100-SXM4-80GB, using Torch 2.9.1+cu126, BF16 autocast and FP32 weights:
| Approximate packed length | Batch 1 requests/s | Fastest tested batch | Requests/s at that batch |
|---|---|---|---|
| 256 | 43.84 | 8 | 273.31 |
| 512 | 41.88 | 8 | 157.26 |
| 2,048 | 13.24 | 8 | 30.67 |
| 4,096 | 4.98 | 8 | 11.44 |
| 8,176 | 1.49 | 4 | 2.96 |
These are throughput measurements, not service-level tail-latency guarantees. Mixed-length batching can be slower than sequential inference. See benchmark methodology.
The current README reports the following approximate timings on one 10-core Apple-silicon laptop using the normal CPU path:
| Reported input length | Approximate time per answer | Approximate peak memory |
|---|---|---|
| A few dozen tokens | 0.05 s | 2 GiB |
| 512 tokens | 0.4 s | 2 GiB |
| 4,096 tokens | 8 s | 5 GiB |
| 8,191 tokens | 30 s | 8 GiB |
These are README usage estimates, not a controlled benchmark or a latency guarantee. The complete packed request must still fit within 8,192 tokens. The README estimates model loading at 5β25 seconds depending on the PyTorch version. Earlier experimental server measurements and RC5 verification timings used different implementations and hardware and should not be treated as measurements of the published 0.1.1 package. In particular, the experimental CPU backend can be slower than the normal backend on Apple silicon; measure on the deployment machine.
Known limitations
- Accepting an 8k input does not guarantee reliable use of all evidence at all positions.
- Option wording, ordering, domain, and input format can affect predictions.
- Training and evaluation are primarily English; broad multilingual quality is not established.
- CPU, GPU, precision, and batching changes can affect logits and occasionally selected labels. Existing GPU calibration must not be assumed to transfer.
- Bias, subgroup fairness, adversarial robustness, and prompt-injection resistance are not comprehensively characterized by the reported benchmark chart.
For deployment, assess representative labeled data, false-positive and false-negative costs, and escalation behavior. Keep deterministic validation around exact arithmetic, schema constraints, and consequential actions when needed.
License and attribution
The BaseDecision software repository is licensed under Apache-2.0. BaseDecision builds on ModernBERT; retain applicable upstream notices. A software repository license alone does not establish the licensing of separately distributed model weights or training datasets. Consult the weight release's license and each dataset's terms before redistribution or deployment.
Sources
- Published Python package, version 0.1.1
- BaseDecision repository and usage
- Third-party evaluation chart
- Calibration profiles and assessment scope
- Measured GPU throughput
Updated October 5, 2026 (America/New_York), using the current README and PyPI 0.1.1 package metadata. Benchmark scores are reproduced as reported.
Model tree for onlyaady/BaseDecision
Base model
answerdotai/ModernBERT-large