Instructions to use edihasaj/surebranch-4b with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use edihasaj/surebranch-4b with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="edihasaj/surebranch-4b")# Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("edihasaj/surebranch-4b") model = AutoModelForCausalLM.from_pretrained("edihasaj/surebranch-4b", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Surebranch 4B
Surebranch 4B scores typed choices from a state in one forward pass. It returns probabilities over the options you provide, without generating an answer string. This is a research release of model weights, tokenizer and configuration. The training code and datasets are not included.
Surebranch 4B is independent of TypeSafe AI and has no training data distilled from Jev.
Use
Install decider-ai>=1.4.0, transformers>=5 and huggingface_hub, then:
from huggingface_hub import snapshot_download
from decider.infer import Decider
model = Decider(snapshot_download("edihasaj/surebranch-4b"), use_graphs=False)
answers = model.decide(
"The customer was charged twice for one order.",
[
{"question": "Which team should handle this?", "options": ["billing", "support", "sales"]},
{"question": "Does this need a refund review?", "options": ["no", "yes"]},
],
)
print(answers)
The decider-ai package supplies the typed request and answer-slot inference interface. This repository contains the merged model weights and the files needed to load them. The confirmed evaluation context was at most 8,192 input tokens. Longer contexts and a production CPU price have not been qualified.
Measured development results
The code set contains 250 twice-executed Python programs, grouped away from this adapter's training programs. Each has a count question, a whole-suite question and ten individual assertions. Ordinary mode asked one hash-selected assertion per program; packed mode asked all twelve questions together.
| Task | Cases | Accuracy |
|---|---|---|
| Ordinary exact passing-test count | 250 | 46.8% |
| Ordinary whole suite passes | 250 | 76.8% |
| Ordinary single assertion passes | 250 | 84.4% |
| Packed exact passing-test count | 250 | 46.8% |
| Packed whole suite passes | 250 | 74.4% |
| Packed single assertion passes | 2,500 | 82.00% |
Packed mode detected 203 of 555 failing assertions. Across 250 separate general decisions from seven families, equal-family macro accuracy was 88.9%. No extra probability calibration is applied beyond the included answer-type temperature configuration.
These are opened development results used for model selection, not an independent final. This checkpoint has not passed a fresh head-to-head Jev parity test. Results are strongest on the tested Python decision format; exact counts and failing assertions remain difficult. Do not treat the probabilities as proof of correctness. Question order can change packed answers, and the model does not execute code.
Origin and license
The fine-tune contributions are provided under CC BY-SA 4.0. The upstream base weights retain their Apache-2.0 terms. Base-model provenance, public training-data sources and their licenses are credited in NOTICE.md. No training examples or evaluation receipts are distributed here.
- Downloads last month
- 14