BaseDecision Model Card

BaseDecision is a compact, text-only decision model for selecting among caller-provided options, answering boolean questions, and assigning discrete ratings. It combines a ModernBERT-large encoder and additional task adaptation. It processes up to 8,192 packed tokens per request and is approximately 0.42 billion parameters.

BaseDecision received approximately 320 million tokens of additional training. In the supplied third-party evaluation, it achieved the highest unweighted average score among the five compared models and led five of eight benchmarks.

Model details

Property Description
Model name BaseDecision
Developer Hrudayaditya β€œAady” Jallu
Repository hrudayaditya/BaseDecision
Model family ModernBERT-large encoder
Approximate size 0.42B parameters; rounded model size, not an exact parameter inventory
Architecture Bidirectional transformer encoder with typed decision heads
Input Text context, question/instructions, and candidate options or rating levels
Output Selected answer, option ID, raw logits, option probabilities, and request metadata
Decision types choice, noul (boolean), and score (discrete rating)
Candidate options 2–255, subject to the total token budget
Context limit 8,192 packed tokens, including context, instructions, options, and special tokens
Primary evaluated language English
Additional training Approximately 320M tokens
Public Python package basedecision; published version reviewed: 0.1.1, Python 3.10+

This is a discriminative decision model. It does not generate free-form explanations, expose a causal language-model .generate() interface, or provide the general-purpose capabilities of a chat model. Multiple questions over one context are evaluated independently; shared transformer-state computation is not claimed.

Intended uses

  • Intent classification and workflow routing with explicit candidate actions.
  • Entity-specific sentiment and stance classification.
  • Textual entailment and document-based decision support.
  • Boolean checks and discrete ratings over supplied text.
  • Local applications that need structured decisions without a remote API call.

Plain-text option labels are supported. Descriptions can help disambiguate labels but are not mandatory. Evaluate the actual question wording, options, input lengths, and data distribution used by the application.

The model should not be treated as an autonomous authority for medical, legal, financial, or security decisions. A correct output schema does not establish factual correctness, safe action selection, or suitability for a high-impact workflow.

Training

Training volume

Approximately 320 million tokens were processed during BaseDecision training, as verified by the developer. This is the additional BaseDecision training volume, not the upstream ModernBERT pretraining corpus size.

Third-party evaluation

The table below reproduces the third-party evaluation results supplied by the developer and published in the repository's benchmark chart. Higher is better. Values are reported benchmark scores; the chart does not identify every row's metric, so they are not uniformly labeled accuracy or F1 here.

Benchmark GLiNER 2.5 base GLiNER2.5-Decide Decision 1.0 Kai 0.6B Laya BaseDecision
BANKING77 23.8 65.6 40.7 14.3 68.2
FinEntity 70.1 66.2 37.0 61.0 71.3
ContractNLI 22.9 21.8 36.6 29.0 66.8
VAST 35.8 35.3 20.8 40.5 66.2
SGD / SGD-X 45.9 0.8 48.5 42.4 51.1
MuSR 34.8 45.2 45.2 43.2 44.6
NLI4CT 39.8 48.0 29.8 47.7 42.3
PhishNChips 50.4 50.0 49.9 50.1 50.0
Unweighted average 40.4 41.6 38.6 41.0 57.6

BaseDecision leads 5 of 8 benchmarks in this comparison: BANKING77, FinEntity, ContractNLI, VAST, and SGD/SGD-X. Its reported average is 16.0 points above the next-highest reported average,

Confidence and calibration

Local inference returns raw softmax probabilities by default. These are relative scores over the supplied choices, not guaranteed probabilities that an answer is correct. Changing the candidate set can change those scores. A high score does not establish evidence completeness or eliminate the possibility of a confidently wrong answer.

Optional scalar-temperature profiles are available for specific assessed workloads:

Profile Release behavior Temperature Assessment NLL, raw β†’ scaled Equal-width ECE, raw β†’ scaled
SGD identifier Explicit opt-in 1.4646 0.6116 β†’ 0.5830 8.05% β†’ 2.97%
SGD schema Explicit opt-in 1.4793 0.6034 β†’ 0.5606 7.88% β†’ 2.96%
VAST Raw recommended; negligible gain 1.0071 0.6246 β†’ 0.6242 7.68% β†’ 7.43%

These are internal calibration assessments, separate from the third-party benchmark table. Each SGD format was assessed on 1,270 questions from 125 dialogue groups; the formats share groups. VAST used 118 questions/32 groups; CLINC used 156 questions/groups. SGD's short and two-option slices regressed despite aggregate gains.

Deployment and usage

Install the published Python package:

pip install basedecision

Version 0.1.1 is published on PyPI. Python 3.10 or newer is required. The default installation includes PyTorch, Transformers, SafeTensors, Tokenizers, Hugging Face Hub, OpenAI, and Anthropic dependencies. The initial download can therefore be large. To reproduce the reviewed package release, use pip install basedecision==0.1.1; pin the runtime and model revision separately when exact reproducibility matters.

The older runtime, hub, and provider extras remain compatibility aliases; they are no longer needed to install those dependencies. Advanced users who already manage their runtime can use pip install --no-deps basedecision and install the dependencies they need themselves. From a repository clone, use pip install ..

Installing the package does not install model weights. Supply a local checkpoint folder. The README's Hugging Face model identifier remains a placeholder; load_from_hub(repository_id, revision=commit_hash) supports downloading a separately published checkpoint, with local_files_only=True available for cached operation. Do not use the placeholder as an actual model ID.

Load an exported checkpoint containing model.safetensors, its encoder configuration, decision configuration, and tokenizer:

from basedecision import load

model = load('/path/to/model')
result = model.choose(
    context='I was charged twice. Please refund the duplicate payment.',
    question='What does the customer request?',
    options=['Refund request', 'Delivery status', 'Change address'],
)
print(result.answer)
print(result.probabilities)  # Raw softmax by default; not calibrated confidence.
model.close()

The current public SDK selects a suitable CUDA path when available and otherwise uses CPU FP32. Explicit device choices are supported. Apple-silicon Macs use the CPU path; native MPS acceleration is not claimed. The opt-in cpu_fast backend uses BF16 and local-attention optimizations, is experimental, checks compatibility at load time, and may be substantially slower on unsuitable hardware. It does not silently substitute another backend. The normal runtime allows Transformers >=4.48,<6, while cpu_fast supports only its checked Transformers 4.48–4.57 range. A default installation can therefore resolve to Transformers 5 and work with the normal backend while cpu_fast raises CPUFastUnavailable. Use a compatible runtime deliberately if testing this experimental path.

Inputs are packed with intact options; over-budget requests raise ContextLengthError rather than being silently truncated. The 8,192-token limit covers the entire packed request, not 8,192 context tokens plus unlimited instructions and options. Simple chunking or voting does not guarantee preservation of cross-document reasoning.

OpenAI and Anthropic adapters provide an alternative SDK backend. They call those providers' models, not BaseDecision weights; their results must not be attributed to BaseDecision. They require explicit selection and send inputs to the chosen provider. Cloud results do not supply BaseDecision's local option probabilities.

Python interfaces and resource management

Interface Purpose
model.choose(...) Select an option from plain strings or Option objects
model.check(...) Return a boolean decision
model.score(...) Select a rating level; optionally return its numeric value and expected value
decide(model, context=..., questions=...) Answer several typed questions independently
model.predict_batch(...) / model.predict_iter(...) Process request collections or streams
result.to_dict() Export a result as ordinary JSON-compatible values
model.count_tokens(request) Inspect the packed request length before inference

Use with load('/path/to/model') as model: or call model.close() to release resources. Closing is idempotent; subsequent prediction on a closed model raises InputError.

The SystemOne adapter accepts the Jev/SystemOne state and typed questions wire format. It is a local-model interface with one forward pass per question; it does not make BaseDecision a Jev model or add shared-context computation. A reference HTTP server is provided in the repository's examples/ directory and binds to localhost. Review SystemOne documentation before exposing it over a network.

Errors and privacy

  • ContextLengthError reports an over-budget input, including required and maximum token counts.
  • InputError reports invalid request types, duplicate options, or other malformed input. Context must be text; serialize structured data explicitly.
  • CPUFastUnavailable reports an unsupported or failed experimental CPU configuration.
  • Cloud failures raise ProviderError with code, retryable, and status_code. The documented error representation excludes API keys and request text.
  • Refused or invalid provider responses raise errors rather than producing a guessed label or silently returning False.

Local inference requires no provider API key. Explicit cloud-backend selection sends text to that provider; there is no automatic cloud fallback. See the API reference, cloud documentation, and security guidance.

Performance and hardware

Latency depends on sequence length, option count, batching, processor instruction support, runtime, and system load. No universal CPU latency guarantee is made.

The repository reports these warmed, synthetic, similar-length throughput measurements on an A100-SXM4-80GB, using Torch 2.9.1+cu126, BF16 autocast and FP32 weights:

Approximate packed length Batch 1 requests/s Fastest tested batch Requests/s at that batch
256 43.84 8 273.31
512 41.88 8 157.26
2,048 13.24 8 30.67
4,096 4.98 8 11.44
8,176 1.49 4 2.96

These are throughput measurements, not service-level tail-latency guarantees. Mixed-length batching can be slower than sequential inference. See benchmark methodology.

The current README reports the following approximate timings on one 10-core Apple-silicon laptop using the normal CPU path:

Reported input length Approximate time per answer Approximate peak memory
A few dozen tokens 0.05 s 2 GiB
512 tokens 0.4 s 2 GiB
4,096 tokens 8 s 5 GiB
8,191 tokens 30 s 8 GiB

These are README usage estimates, not a controlled benchmark or a latency guarantee. The complete packed request must still fit within 8,192 tokens. The README estimates model loading at 5–25 seconds depending on the PyTorch version. Earlier experimental server measurements and RC5 verification timings used different implementations and hardware and should not be treated as measurements of the published 0.1.1 package. In particular, the experimental CPU backend can be slower than the normal backend on Apple silicon; measure on the deployment machine.

Known limitations

  • Accepting an 8k input does not guarantee reliable use of all evidence at all positions.
  • Option wording, ordering, domain, and input format can affect predictions.
  • Training and evaluation are primarily English; broad multilingual quality is not established.
  • CPU, GPU, precision, and batching changes can affect logits and occasionally selected labels. Existing GPU calibration must not be assumed to transfer.
  • Bias, subgroup fairness, adversarial robustness, and prompt-injection resistance are not comprehensively characterized by the reported benchmark chart.

For deployment, assess representative labeled data, false-positive and false-negative costs, and escalation behavior. Keep deterministic validation around exact arithmetic, schema constraints, and consequential actions when needed.

License and attribution

The BaseDecision software repository is licensed under Apache-2.0. BaseDecision builds on ModernBERT; retain applicable upstream notices. A software repository license alone does not establish the licensing of separately distributed model weights or training datasets. Consult the weight release's license and each dataset's terms before redistribution or deployment.

Sources

Updated October 5, 2026 (America/New_York), using the current README and PyPI 0.1.1 package metadata. Benchmark scores are reproduced as reported.

Downloads last month

-

Downloads are not tracked for this model. How to track
Safetensors
Model size
0.4B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for onlyaady/BaseDecision

Finetuned
(389)
this model