Instructions to use morriszjm/Tacit-4B with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use morriszjm/Tacit-4B with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-generation", model="morriszjm/Tacit-4B") messages = [ {"role": "user", "content": "Who are you?"}, ] pipe(messages)# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModelForCausalLM tokenizer = AutoTokenizer.from_pretrained("morriszjm/Tacit-4B") model = AutoModelForCausalLM.from_pretrained("morriszjm/Tacit-4B", device_map="auto") messages = [ {"role": "user", "content": "Who are you?"}, ] inputs = tokenizer.apply_chat_template( messages, add_generation_prompt=True, tokenize=True, return_dict=True, return_tensors="pt", ).to(model.device) outputs = model.generate(**inputs, max_new_tokens=256) print(tokenizer.decode(outputs[0][inputs["input_ids"].shape[-1]:])) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- vLLM
How to use morriszjm/Tacit-4B with vLLM:
Install from pip and serve model
# Install vLLM from pip: pip install vllm # Start the vLLM server: vllm serve "morriszjm/Tacit-4B" # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "morriszjm/Tacit-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker
docker model run hf.co/morriszjm/Tacit-4B
- SGLang
How to use morriszjm/Tacit-4B with SGLang:
Install from pip and serve model
# Install SGLang from pip: pip install sglang # Start the SGLang server: python3 -m sglang.launch_server \ --model-path "morriszjm/Tacit-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "morriszjm/Tacit-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }'Use Docker images
docker run --gpus all \ --shm-size 32g \ -p 30000:30000 \ -v ~/.cache/huggingface:/root/.cache/huggingface \ --env "HF_TOKEN=<secret>" \ --ipc=host \ lmsysorg/sglang:latest \ python3 -m sglang.launch_server \ --model-path "morriszjm/Tacit-4B" \ --host 0.0.0.0 \ --port 30000 # Call the server using curl (OpenAI-compatible API): curl -X POST "http://localhost:30000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "morriszjm/Tacit-4B", "messages": [ { "role": "user", "content": "What is the capital of France?" } ] }' - Docker Model Runner
How to use morriszjm/Tacit-4B with Docker Model Runner:
docker model run hf.co/morriszjm/Tacit-4B
Tacit-4B (preview)
Tacit-4B makes typed decisions (choose an option, answer yes/no, or pick an ordered level) in a single forward pass, returning a probability over the options. It is built on Qwen/Qwen3-4B with AnyJev self-distillation: the base model wrote its own decision problems, answered them with its reasoning on, and learned to give those answers in one forward pass. No human labels, external datasets or other models were used.
Results
Accuracy in one forward pass on the full test sets, with the same prompt and readout for both models.
| benchmark | Qwen3-4B | Tacit-4B | difference |
|---|---|---|---|
| JevBench (public set, 231) | 0.658 | 0.723 | +6.5 pts |
| bev-decision (test split, 46,320) | 0.619 | 0.663 | +4.4 pts |
Usage
from transformers.dynamic_module_utils import get_class_from_dynamic_module
Tacit = get_class_from_dynamic_module("tacit_decision.Tacit", "morriszjm/Tacit-4B")
# one forward per decision
tacit = Tacit.from_pretrained("morriszjm/Tacit-4B")
# or: send low-confidence decisions to the model's own reasoning
tacit = Tacit.from_pretrained("morriszjm/Tacit-4B", adaptive=True, tau=0.5, max_cot_share=0.2)
# or: the same on vLLM (pip install vllm), for throughput
tacit = Tacit.from_pretrained("morriszjm/Tacit-4B", engine="vllm", adaptive=True)
d = tacit.decide(state="Customer: my package was due last Monday and it still has not arrived.",
question="What does the customer want?",
options=["track_order", "cancel_order", "refund", "change_address"])
print(d["answer"], d["probs"], d["route"]) # route: "one_forward" or "cot"
kind is "choice" (default), "yes_no", or "score" (options are ordered levels, lowest
first); decide_batch([...]) takes a list of such dicts.
| argument | default | meaning |
|---|---|---|
adaptive |
False |
send low-confidence decisions to the model's own reasoning (thinking on) |
tau |
0.5 |
a decision is low-confidence when the log-probability gap between its top two options is below tau |
max_cot_share |
0.2 |
at most this share of the last cot_window decisions goes to reasoning; None removes the cap |
cot_window |
1000 |
how many recent decisions the cap counts; None counts every decision since loading |
cot_max_tokens |
8192 |
reasoning budget of one decision |
engine |
"transformers" |
"vllm" runs vLLM in this process; "server" uses a running vllm serve (see below) |
base_url |
None |
engine="server": the server's /v1 URL |
vllm_kwargs |
None |
passed to vllm.LLM, e.g. dict(gpu_memory_utilization=0.85, max_model_len=32768) |
An escalated decision's answer is read from the label distribution after the reasoning, so it also
comes with probabilities. Without adaptive, nothing is generated. On vLLM the labels are read from
the top 20 log-probabilities of the full vocabulary (a label outside them counts as 0). The checkpoint also loads as a
plain causal LM (AutoModelForCausalLM, vLLM); tacit_decision.py holds the prompt it was trained
with.
Serving with vLLM
vllm serve morriszjm/Tacit-4B --host 127.0.0.1 --port 8000
tacit = Tacit.from_pretrained("morriszjm/Tacit-4B", engine="server",
base_url="http://127.0.0.1:8000/v1", adaptive=True)
Only the tokenizer loads on the client. For many clients sharing one cap, run the HTTP gateway that ships in this repo in front of the server:
hf download morriszjm/Tacit-4B tacit_decision.py --local-dir .
python tacit_decision.py serve --model morriszjm/Tacit-4B --upstream http://127.0.0.1:8000/v1 --adaptive --port 8100
curl -s 127.0.0.1:8100/v1/decide -H 'Content-Type: application/json' -d '{"state": "Customer: my package has not arrived.", "question": "What does the customer want?", "options": ["track_order", "refund"]}'
Apache-2.0, as is the base model by the Qwen team.
- Downloads last month
- 928