ZooWork Instinct Tuned 4B

GitHub Hosted API JevBench submission License

ZooWork · Instinct homepage and hosted API · API model ID: instinct-tuned-4b

Instinct Tuned 4B is a 4B-parameter decision model. You give it a shared state (a message, a document, or a JSON object) and a question with a fixed set of candidate answers. It returns a calibrated probability for each candidate from one forward pass. It does not generate text.

Do not use generate() or a transformers pipeline with this model. The decision is read from the label-token logits with a fixed prompt and temperature; use the reference runtime below.

It supports three question types:

Type Question Output
choice Pick one of 2–16 described options a probability per option, the argmax, and a confidence
noul Is this proposition true of the state? P(yes)
score Place the state on an ordered scale of 2–16 levels a probability per level, and the expected level

The model is a LoRA fine-tune, merged into the base weights, of Qwen/Qwen3.5-4B.

Quick start

The readout is not a standard generate() or classification head, so use the reference runtime: SerendipityOneInc/instinct.

git clone https://github.com/SerendipityOneInc/instinct && cd instinct
pip install -e ".[gpu]"
instinct-decide --model instinct-tuned-4b models/instinct-tuned-4b/examples/request.json
from instinct import InstinctModel, answer

model = InstinctModel.from_pretrained("instinct-tuned-4b")  # downloads srpone/instinct-tuned-4b
print(answer(model, {
    "state": "Where is my package? I ordered it last week and it still hasn't arrived.",
    "questions": {
        "intent": {"type": "choice", "instructions": "Which intent does the message express?",
                   "criteria": {"track_order": "Wants to know where an order is",
                                "cancel_order": "Wants to cancel an order",
                                "billing_question": "Asks about a charge or payment"}},
        "angry": {"type": "noul", "instructions": "The customer is angry."},
    },
}))

How it works

  1. The request is rendered with the fixed prompt instinct.prompt.v1:
    • Shared state:\n<state>\n\n
    • then a sorted-key JSON object {"criteria": [{"description", "label"}], "instructions", "primitive"}, with the candidates labelled A, B, …
    • then Return only the selected letter: A, B, ….\nAnswer:
  2. The prompt goes through the chat template with add_generation_prompt=True and enable_thinking=False.
  3. The last hidden state at full depth (layer 32) passes through the final norm. Only the LM-head rows of the candidate label tokens are applied ("candidate-row readout").
  4. The logits are divided by the serving temperature T = 2.80, then softmaxed. T is stored in decision_config.json.

A noul question is scored as a two-way choice between the canonical candidates yes ("The stated proposition is true.") and no. Custom true/false wording is accepted but not shown to the model. A score question is scored as a choice over levels "0"…"n-1". A structured state is rendered with json.dumps(state, ensure_ascii=False).

Evaluation

The table reports JevBench public (231 items; 198/231 overall), full depth, bf16, scored with the official JevBench client, items in their original option order:

Scope Correct Accuracy
All published tasks 198/231 85.71%
easy 48/48 100.00%
standard 69/72 95.83%
hard 81/111 72.97%

These public items were also used during our development for model selection, so treat the table as a reference point, not a held-out score.

Calibration: the expected calibration error on the hard split after temperature scaling is 0.04–0.11, depending on option order (see Limitations).

Serving latency Value
Direct-upstream p50 62 ms
Direct-upstream p95 ≈110 ms (estimated)

Latency refers to warmed, serial serving and excludes public Internet, TLS and gateway overhead. The p50 is measured; the p95 is an estimate from the measured direct-upstream p50 and observed end-to-end tail spread, not an SLO.

Training

  • Base: Qwen3.5-4B, text path only. The vision tower weights are carried over from the base unchanged.
  • Method: LoRA (r = 32, α = 64) on the decision objective (cross-entropy over the candidate label tokens), with auxiliary readouts at layers 16 and 24. The auxiliary readouts are not used at inference. Settings: 1 epoch, learning rate 1e-4, effective batch 32. The LoRA was merged in fp32 and saved in bf16.
  • Data: about 9.4k decision items: a base decision training set (~5.9k), plus ~900 synthetic hard examples blind-verified by an independent model and upsampled 4x. The training data is not released at this time; we plan to release it.
  • Decontamination (training data only): every training item has less than 20% 13-gram overlap with the JevBench public set and with our held-out development sets. This check covers the training data; the public set itself was used for model selection (see Evaluation).
  • Temperature: T was fitted on a held-out development set drawn from the same task distribution as the benchmark.

Limitations

  • Option-order sensitivity. Probabilities depend on which candidate gets which letter. Accuracy is stable, but calibration error on hard items moves by 0.04–0.07 between orders. The runtime uses the order given in the request.
  • Text only, 8,192-token limit. Longer inputs are rejected, not truncated. Image and video inputs are not supported.
  • Mostly English evaluation.
  • T was fitted on a held-out development set from the same task distribution as the benchmark, so the probabilities are calibrated for that distribution. Re-fit T on your own labelled data if your domain differs.
  • The model makes decisions; it does not explain them.

Files

File Purpose
model-0000{1,2,3}-of-00003.safetensors bf16 weights
model.safetensors.index.json shard index
config.json, generation_config.json Qwen3.5 config. The top-level dtype is bfloat16; text_config and vision_config say float32 because that was the merge precision. The shards are bf16.
tokenizer.json, tokenizer_config.json, chat_template.jinja tokenizer and chat template
decision_config.json prompt version, serving temperature, readout depth
SHA256SUMS file hashes

License

The weights are released under Apache-2.0, the license of the base model Qwen3.5-4B.

About

Instinct is developed by ZooWork. Learn more at instinct.zoowork.ai.

Downloads last month
-
Safetensors
Model size
5B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for srpone/instinct-tuned-4b

Finetuned
Qwen/Qwen3.5-4B
Finetuned
(801)
this model

Collection including srpone/instinct-tuned-4b

Article mentioning srpone/instinct-tuned-4b