Tacit-9B (preview)

Tacit-9B makes typed decisions (choose an option, answer yes/no, or pick an ordered level) in a single forward pass, returning a probability over the options. It is built on Qwen/Qwen3.5-9B with AnyJev self-distillation: the base model wrote its own decision problems, answered them with its reasoning on, and learned to give those answers in one forward pass. No human labels, external datasets or other models were used.

Results

Accuracy in one forward pass on the full test sets, with the same prompt and readout for both models.

benchmark Qwen3.5-9B Tacit-9B difference
JevBench (public set, 231) 0.805 0.823 +1.7 pts
bev-decision (test split, 46,320) 0.680 0.725 +4.5 pts

Usage

from transformers.dynamic_module_utils import get_class_from_dynamic_module

Tacit = get_class_from_dynamic_module("tacit_decision.Tacit", "morriszjm/Tacit-9B")
# one forward per decision
tacit = Tacit.from_pretrained("morriszjm/Tacit-9B")
# or: send low-confidence decisions to the model's own reasoning
tacit = Tacit.from_pretrained("morriszjm/Tacit-9B", adaptive=True, tau=0.5, max_cot_share=0.2)
# or: the same on vLLM (pip install vllm), for throughput
tacit = Tacit.from_pretrained("morriszjm/Tacit-9B", engine="vllm", adaptive=True)

d = tacit.decide(state="Customer: my package was due last Monday and it still has not arrived.",
                 question="What does the customer want?",
                 options=["track_order", "cancel_order", "refund", "change_address"])
print(d["answer"], d["probs"], d["route"])  # route: "one_forward" or "cot"

kind is "choice" (default), "yes_no", or "score" (options are ordered levels, lowest first); decide_batch([...]) takes a list of such dicts.

argument default meaning
adaptive False send low-confidence decisions to the model's own reasoning (thinking on)
tau 0.5 a decision is low-confidence when the log-probability gap between its top two options is below tau
max_cot_share 0.2 at most this share of the last cot_window decisions goes to reasoning; None removes the cap
cot_window 1000 how many recent decisions the cap counts; None counts every decision since loading
cot_max_tokens 8192 reasoning budget of one decision
engine "transformers" "vllm" runs vLLM in this process; "server" uses a running vllm serve (see below)
base_url None engine="server": the server's /v1 URL
vllm_kwargs None passed to vllm.LLM, e.g. dict(gpu_memory_utilization=0.85, max_model_len=32768)

An escalated decision's answer is read from the label distribution after the reasoning, so it also comes with probabilities. Without adaptive, nothing is generated. On vLLM the labels are read from the top 20 log-probabilities of the full vocabulary (a label outside them counts as 0). The checkpoint also loads as a plain causal LM (AutoModelForCausalLM, vLLM); tacit_decision.py holds the prompt it was trained with.

Serving with vLLM

vllm serve morriszjm/Tacit-9B --host 127.0.0.1 --port 8000
tacit = Tacit.from_pretrained("morriszjm/Tacit-9B", engine="server",
                              base_url="http://127.0.0.1:8000/v1", adaptive=True)

Only the tokenizer loads on the client. For many clients sharing one cap, run the HTTP gateway that ships in this repo in front of the server:

hf download morriszjm/Tacit-9B tacit_decision.py --local-dir .
python tacit_decision.py serve --model morriszjm/Tacit-9B --upstream http://127.0.0.1:8000/v1 --adaptive --port 8100
curl -s 127.0.0.1:8100/v1/decide -H 'Content-Type: application/json' -d '{"state": "Customer: my package has not arrived.", "question": "What does the customer want?", "options": ["track_order", "refund"]}'

On Hopper GPUs vLLM compiles this model's gated-delta-rule kernel with nvcc on first use, so the serving machine needs the CUDA toolkit (engine="vllm" in Python switches to a Triton kernel by itself).

Apache-2.0, as is the base model by the Qwen team.

Downloads last month
132
Safetensors
Model size
9B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for morriszjm/Tacit-9B

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(982)
this model

Collection including morriszjm/Tacit-9B