Sors-2B

Sors-2B is a 2B decision model that turns a state and a set of typed questions into decisions. It reads the state as text or JSON and returns a probability for every allowed option of every question. There is no free-form text generation and no output parsing.

The Sors API follows the Jev / SystemOne request and response format.

Sors-2B is fine-tuned from Qwen/Qwen3.5-2B-Base.

Model

SORS architecture

  • Backbone: Qwen/Qwen3.5-2B-Base text backbone, fully fine-tuned.
  • Decision tokens: every option row on the menu starts with one of 256 decision tokens (<|D0|> … <|D255|>). The tokens carry no meaning of their own: during training each question gets a fresh, balanced set of codes, so the model can only answer by reading the option text. Options are defined per request, up to 255 per question, with no retraining.
  • Decision layers: attached to the backbone's last layers. Option rows read the state and question, compare with each other without any position encoding between options, and write option-set information back into the backbone. A shared scorer gives one score per option.
  • Output: one probability per allowed option, normalised over that question's options only.

The backbone keeps the causal order of state → question → options and the decision layers remove the order among options. Training enforces the rest: every question is shown twice in two random option orders and a Jensen–Shannon term makes the two answers agree, codes are re-drawn per question, half the questions hide option descriptions, and half the menus add a "none of these" option.

Training data

Sors-2B is trained only on our own synthetic dataset, synth-intents-v5.3 (to be open-sourced once it is cleaned up). None of the evaluation sets below take part in training. The synthetic data is written from scratch, contains no content from existing datasets, and stays out of Banking77's banking domain and MASSIVE's voice-assistant scenarios. Compared against every evaluation set below, no text is identical after normalisation, and only 4 of Banking77's 3,080 items share a run of 8 consecutive words with the training material, all of them everyday phrases.

Files

File Purpose
model.safetensors All weights (3.89 GB): backbone in BF16, decision layers in FP32, learned decision-token rows
config.json Sors architecture, text-model configuration, decision-token IDs and per-parameter precision
tokenizer.json, tokenizer_config.json, chat_template.jinja Training tokenizer, including its 261 added special tokens
LICENSE Apache-2.0 license

Usage

Sors-2B runs with the SORS runtime from the code repository. Tested with torch 2.14 and transformers 5.17 on a single A100 40GB.

git clone https://github.com/madousho-ai/sors && cd sors
uv sync
PYTHONPATH=src .venv/bin/python scripts/serve.py \
  --init SakuraYuyuko/Sors-2B --warmup --port 8000

--init downloads this repository into the standard Hugging Face cache on first start.

Jev / SystemOne API

The server answers POST /v1/systemone. The response has model, answers keyed by question ID, and usage. A choice answer has choice, confidence and probabilities; a score answer has the expected score, confidence, legend and probabilities; a noul answer has the probability of true.

curl -s http://127.0.0.1:8000/v1/systemone \
  -H 'Content-Type: application/json' \
  -d '{
    "model": "Sors-2B",
    "state": "I need to cancel my subscription before it renews.",
    "questions": {
      "intent": {
        "type": "choice",
        "instructions": "What does the customer want?",
        "criteria": {"Cancel a subscription": null, "Reset a password": null, "Track a shipment": null}
      },
      "urgent": {"type": "noul", "instructions": "Does the customer need this done right away?"}
    }
  }'

Python

from pathlib import Path

from huggingface_hub import snapshot_download
from sors.serve.api import Choice
from sors.serve.engine import load_engine

engine = load_engine(Path(snapshot_download("SakuraYuyuko/Sors-2B")), device="cuda", local_files_only=True)
question = Choice(
    type="choice",
    instructions="What does the customer want?",
    criteria={"Cancel a subscription": None, "Reset a password": None, "Track a shipment": None},
)
result = engine.evaluate("I need to cancel my subscription before it renews.", {"intent": question})
print(dict(zip(question.criteria, result.probs["intent"])))

Input format

Field Description
model Served model name, Sors-2B by default
state Any string or JSON value describing the situation to decide on
questions Mapping of question ID to question; several questions share one state

Each question has:

  • type: noul (yes/no), choice (named options), or score (ordered options)
  • instructions: what to decide, as a string or JSON
  • criteria: for choice, a mapping of option name to an optional description (2–255 options); for score, a list of ordered levels; for noul, optional descriptions for true and false

Inputs are limited to 8,192 tokens. score is supported by the API but has no dedicated training data yet.

Results

Accuracy (%). Banking77 and MASSIVE are answered against the full intent menu: all 77 and all 60 intents on every question, with no top-K shortlist. "+ descriptions" adds a one-line description after each intent name.

Benchmark Questions Sors-2B
Banking77 3,080 70.3
Banking77 + descriptions 3,080 81.1
MASSIVE 2,974 72.3
MASSIVE + descriptions 2,974 81.6
BoolQ 3,270 84.0
Public JevBench easy 48 100.0
Public JevBench original 72 90.3
Public JevBench hard 111 73.0
Public JevBench overall 231 84.0
  • 17 of Banking77's intents never appear in training.
  • With options shuffled into 5 random orders, accuracy stays within 1.3 points of the table, and the 5 orders pick the same option 0.86 of the time on Banking77, 0.82 on MASSIVE and 0.98 on BoolQ.
  • Public JevBench is JevBench's public question set. Its tiers are small, so differences of a few points there are noise. Results from the JevBench leaderboard will be added once available.

License

Model weights are released under Apache-2.0, following the Qwen base model. The API format follows TypeSafe's System One; this project is not affiliated with TypeSafe AI.

Downloads last month
10
Safetensors
Model size
2B params
Tensor type
F32
·
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for SakuraYuyuko/Sors-2B

Finetuned
(105)
this model