StartLux-Decision-4B

StartLux-Decision: a probability for every option

StartLux-Decision-4B answers typed questions about a state: pick one of several options, yes or no, or a rating on a scale. Every question comes back with a probability for each option. Requests and responses use the TypeSafe /v1/systemone format, so clients written for Jev work unchanged.

Sizes: 0.8B · 2B · 4B · 9B · 27B

Results

StartLux-Decision-4B
Decision Index 0.2.1 52.75
Decision Index 0.2 48.38
JevBench public, correct of 231 204
Intern-Decision, average accuracy over seven suites 91.17
Latency, one request with three questions 26.0 ms
Latency, one yes/no question 14.7 ms

Latency is end to end over HTTP on one H200 in bf16, one request at a time.

Decision Index 0.2.1

Decision Index at every size

Compared with other decision models

Model JevBench public, of 231 Intern avg DI 0.2 / 0.2.1 Latency, 3 questions
StartLux-Decision-27B 208 91.82 59.54 / 63.88 102.3 ms
StartLux-Decision-9B 201 91.08 54.37 / 58.63 35.7 ms
StartLux-Decision-4B 204 91.17 48.38 / 52.75 26.0 ms
Intern-Decision-4B 201 90.02 35.90 / 37.81 44.2 ms ¹
JevK5 200 85.16 36.44 / 38.81
Jev 1.13 199 88.74 51.67 / 57.91 64.0 ms ²
StartLux-Decision-2B 196 88.46 40.72 / 44.19 15.5 ms
SemIf 187 84.23 25.70 / 25.94
Intern-Decision-2B 180 84.68 19.49 / 19.38 33.3 ms ¹
StartLux-Decision-0.8B 179 85.03 35.57 / 38.86 12.2 ms
Intern-Decision-0.8B 163 79.38 11.32 / 11.94 34.0 ms ¹
Laya 130 57.77 5.51 / 6.04

JevBench public counts the correct answers on the 231 public items in the Intern-Decision bundle; Intern avg is the average accuracy over its seven suites; DI is the Decision Index under both editions, the public board's values for the other systems. A blank cell means the number is not published. StartLux-Decision latencies are for one H200, with the three questions answered in one forward pass. ¹ Intern-Decision's own measurement on an RTX 4090. ² The server time the TypeSafe API gateway reports for the same request, mean of 100, network left out as in ours; Intern-Decision reports 109.7 ms end to end.

Latency on one H200

Fast inference

The folder ships its own inference package, startlux_decision/, which is the fast path:

  • requirements.txt installs the fast kernels, flash-linear-attention and causal-conv1d, and python -m startlux_decision.check . confirms they are active. Without them transformers falls back to a path more than ten times slower, and the server refuses to start on a GPU.
  • All questions of a request run in one forward pass, and the server records CUDA graphs at start-up and replays them for short requests: on one H200 a request with three questions takes 26.0 ms end to end and a single yes/no question 14.7 ms.
  • For bulk work, decide_batch batches the questions of many requests together, which is several times faster than sending them one at a time.

Usage

hf download startlux-models/StartLux-Decision-4B --local-dir StartLux-Decision-4B
cd StartLux-Decision-4B
pip install -r requirements.txt
python -m startlux_decision.check .                      # must print "fast kernels: active"
python -m startlux_decision.server --model . --port 8090
curl -s localhost:8090/v1/systemone -H 'Content-Type: application/json' -d '{
  "state": {"ticket": "I was charged twice for order #4411 and the app still shows it as unpaid."},
  "questions": {
    "team":   {"type": "choice", "instructions": "Which team should handle this ticket?",
               "criteria": {"billing": "Payments, refunds and invoices",
                            "shipping": "Delivery and tracking",
                            "technical": "App, login and account problems"}},
    "urgent": {"type": "noul", "instructions": "Should this ticket be answered today?"},
    "severity": {"type": "score", "instructions": "How severe is the impact?",
                 "criteria": ["cosmetic", "annoying", "blocks the customer"]}
  }
}'

Or in Python, from the same folder:

from startlux_decision import StartLuxDecision

m = StartLuxDecision(".")
answers, usage = m.decide(state, questions)
many = m.decide_batch([(state, questions), ...])

License

The model weights are released under CC BY-NC 4.0: free for research and other non-commercial use, with attribution. Commercial use requires a separate license from StartLux Labs; contact contact@startlux.com. The inference code in startlux_decision/ is Apache-2.0. See LICENSE and NOTICE.

Downloads last month
2
Safetensors
Model size
5B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for startlux-models/StartLux-Decision-4B

Quantizations
4 models

Collection including startlux-models/StartLux-Decision-4B