Bespoke-Nimble-9B for vllm.cpp
This repository holds bespokelabs/Bespoke-Nimble-9B
merged into its base Qwen/Qwen3.5-9B
and laid out as the NimbleModel directory that
vllm.cpp loads.
Nimble is a decision model from Bespoke Labs. For each question it runs one
forward pass and reads the logits of the answer letters at the last prompt
position. It does not generate text. vllm.cpp serves it on /v1/systemone
for choice, noul and score questions.
This checkpoint only works with vllm.cpp (directly, or through LocalAI's
vllm-cpp backend). The architecture name is NimbleModel, which
transformers and vLLM do not know. To run Nimble with the authors' own code,
use the upstream adapter.
Redistribution and license
This is a redistribution of upstream weights, converted for vllm.cpp and LocalAI. No weights were trained here. The only change is the LoRA merge described below.
- Adapter, tokenizer, chat template, schema and temperature config:
bespokelabs/Bespoke-Nimble-9B
@
bd792f44ec8e265be861bfcdf4e05967ffe0e858, by Bespoke Labs, Apache-2.0. Project: bespokelabsai/nimble. - Base weights: Qwen/Qwen3.5-9B
@
c202236235762e1c871ad0ccb60c8ee5ba337b9a, by the Qwen team, Apache-2.0.
The upstream licenses apply to these files. Credit for the model belongs to
the original authors. This is the checkpoint without a suffix (temperature
1.0). Bespoke-Nimble-9B-v2 is a different, older checkpoint and is not in
this repository.
Conversion
The converter is scripts/convert-nimble.py from vllm.cpp. The directory was
produced on 2026-09-30 with that script as it is at vllm.cpp commit
96788348627b6a079fcc3ef7fc6676a970965d7b (the script did not change between
the conversion run and that commit). It checks the adapter's SHA256SUMS and
prompt-contract hash, pairs all 248 LoRA modules with base tensors of the
right shape, and merges them the PEFT way:
W' = bf16(float32(W) + (B @ A) * 2) (r=16, alpha=32).
hf download Qwen/Qwen3.5-9B --revision c202236235762e1c871ad0ccb60c8ee5ba337b9a \
--local-dir Qwen3.5-9B
hf download bespokelabs/Bespoke-Nimble-9B \
--revision bd792f44ec8e265be861bfcdf4e05967ffe0e858 --local-dir Bespoke-Nimble-9B
python3 scripts/convert-nimble.py Bespoke-Nimble-9B \
--base-model-dir Qwen3.5-9B --output-dir nimble-9b
config.json records the provenance: adapter sha256
29ef39b072dee97287947455337879c1e916705c2f727287922a2d81f5e2f20a, base
revision, task schema_candidate_classification_v2, and the list of files
checked against SHA256SUMS.
Files
| file | bytes | content |
|---|---|---|
model.safetensors-00001-of-00004.safetensors |
5,276,436,248 | merged weights, BF16 (some F32 tensors as in the base) |
model.safetensors-00002-of-00004.safetensors |
5,335,161,536 | merged weights |
model.safetensors-00003-of-00004.safetensors |
5,368,717,472 | merged weights |
model.safetensors-00004-of-00004.safetensors |
3,325,995,744 | merged weights |
model.safetensors.index.json |
79,657 | shard index, same 775 tensor names as the base |
config.json |
3,463 | base config, architectures: ["NimbleModel"], nimble_temperature: 1.0, nimble_max_length: 8192, provenance |
tokenizer.json, tokenizer_config.json, chat_template.jinja |
from the adapter repo | |
schema_config.json, temperature_config.json, adapter_config.json |
from the adapter repo, kept for provenance |
The shards keep the base's shard names and full tensor set. Total weight size is about 19.3 GB.
Serve it with vllm.cpp
Build the vllm.cpp server (-DVLLM_CPP_SERVER=ON, target vllm-server), then:
hf download mudler/Bespoke-Nimble-9B-vllm-cpp --revision 52eead25b6723710ec16378942c9ae914d179b6f --local-dir nimble-9b
build/examples/vllm-server --model nimble-9b --served-model-name nimble --port 8000
Serve it with LocalAI
A model config for LocalAI's vllm-cpp backend:
name: nimble-9b
backend: vllm-cpp
known_usecases:
- decisions
parameters:
model: mudler/Bespoke-Nimble-9B-vllm-cpp
artifacts:
- name: model
target: model
source:
type: huggingface
repo: mudler/Bespoke-Nimble-9B-vllm-cpp
revision: 52eead25b6723710ec16378942c9ae914d179b6f
Then send the same request body to LocalAI's POST /v1/systemone with
"model": "nimble-9b".
Example
Request (served on CPU from this exact directory, 2026-09-30):
curl http://localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
"model": "nimble",
"state": "I was charged twice for my subscription this month. Please refund the duplicate payment.",
"questions": {
"refund": {"type": "noul", "instructions": "Does the user request a refund?"},
"department": {"type": "choice", "instructions": "Which department should handle this?",
"criteria": {"billing": "Payments and refunds", "technical": "Software bugs", "sales": "New purchases"}},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["Routine", "Urgent", "Emergency"]}}}'
Response:
{
"model": "nimble",
"answers": {
"refund": {"type": "noul", "noul": 0.9966597833278789},
"department": {"type": "choice", "choice": "billing",
"probabilities": {"billing": 0.9881318797907442, "technical": 0.010937236221712353, "sales": 0.0009308839875433919},
"confidence": 0.9383928321620487},
"urgency": {"type": "score", "score": 0.3872836981253824,
"legend": {"0": "Routine", "1": "Urgent", "2": "Emergency"},
"probabilities": {"0": 0.6300001903087783, "1": 0.3527159212570611, "2": 0.017283888434160653},
"confidence": 0.33663361186646856}
},
"usage": {"input_tokens": 988, "output_tokens": 0},
"latency_ms": 75686.59
}
The answer semantics are those of Nimble's reference server (openjev): noul
is p(true), choice is the most probable key, score is the expected level
index, confidence is 1 - H(p) / ln(K), and probabilities are not rounded.
The latency above is a 20-thread CPU and says nothing about GPU speed.
What was verified and what was not
Verified:
- The one live request above, served by the vllm.cpp CPU server built at
commit
96788348627b6a079fcc3ef7fc6676a970965d7b, returns the answers shown. They are plausible for the input. - Earlier on 2026-09-30, the vllm.cpp project compared this same converted
directory, served on
/v1/systemoneon CPU, against the Bespoke authors' ownparallel_schema.prepare_prompts,inference.candidate_logitsandinference.decision_resultovertransformers5.3.0 BF16 with the adapter unmerged through PEFT 0.21.0, also on CPU. Five requests, seven questions (choice, noul with and without criteria, score): argmax equal on 7 of 7, input token counts equal on 5 of 5, largest probability difference 0.0054. That comparison was not repeated for this upload.
Not verified:
- No GPU (CUDA, ROCm, Metal) run. The authors' reference runner requires a CUDA BF16 GPU; the comparison above called its functions on CPU.
- Seven questions are a check, not an accuracy measurement.
- No vLLM gate exists, because vLLM does not serve
NimbleModeldecisions. - vllm.cpp refuses fields with more than 26 choices; the upstream release supports up to 255.
- No GGUF or quantized variant.
- Downloads last month
- 2