Bespoke-Nimble-9B for vllm.cpp

This repository holds bespokelabs/Bespoke-Nimble-9B merged into its base Qwen/Qwen3.5-9B and laid out as the NimbleModel directory that vllm.cpp loads.

Nimble is a decision model from Bespoke Labs. For each question it runs one forward pass and reads the logits of the answer letters at the last prompt position. It does not generate text. vllm.cpp serves it on /v1/systemone for choice, noul and score questions.

This checkpoint only works with vllm.cpp (directly, or through LocalAI's vllm-cpp backend). The architecture name is NimbleModel, which transformers and vLLM do not know. To run Nimble with the authors' own code, use the upstream adapter.

Redistribution and license

This is a redistribution of upstream weights, converted for vllm.cpp and LocalAI. No weights were trained here. The only change is the LoRA merge described below.

The upstream licenses apply to these files. Credit for the model belongs to the original authors. This is the checkpoint without a suffix (temperature 1.0). Bespoke-Nimble-9B-v2 is a different, older checkpoint and is not in this repository.

Conversion

The converter is scripts/convert-nimble.py from vllm.cpp. The directory was produced on 2026-09-30 with that script as it is at vllm.cpp commit 96788348627b6a079fcc3ef7fc6676a970965d7b (the script did not change between the conversion run and that commit). It checks the adapter's SHA256SUMS and prompt-contract hash, pairs all 248 LoRA modules with base tensors of the right shape, and merges them the PEFT way: W' = bf16(float32(W) + (B @ A) * 2) (r=16, alpha=32).

hf download Qwen/Qwen3.5-9B --revision c202236235762e1c871ad0ccb60c8ee5ba337b9a \
  --local-dir Qwen3.5-9B
hf download bespokelabs/Bespoke-Nimble-9B \
  --revision bd792f44ec8e265be861bfcdf4e05967ffe0e858 --local-dir Bespoke-Nimble-9B
python3 scripts/convert-nimble.py Bespoke-Nimble-9B \
  --base-model-dir Qwen3.5-9B --output-dir nimble-9b

config.json records the provenance: adapter sha256 29ef39b072dee97287947455337879c1e916705c2f727287922a2d81f5e2f20a, base revision, task schema_candidate_classification_v2, and the list of files checked against SHA256SUMS.

Files

file bytes content
model.safetensors-00001-of-00004.safetensors 5,276,436,248 merged weights, BF16 (some F32 tensors as in the base)
model.safetensors-00002-of-00004.safetensors 5,335,161,536 merged weights
model.safetensors-00003-of-00004.safetensors 5,368,717,472 merged weights
model.safetensors-00004-of-00004.safetensors 3,325,995,744 merged weights
model.safetensors.index.json 79,657 shard index, same 775 tensor names as the base
config.json 3,463 base config, architectures: ["NimbleModel"], nimble_temperature: 1.0, nimble_max_length: 8192, provenance
tokenizer.json, tokenizer_config.json, chat_template.jinja from the adapter repo
schema_config.json, temperature_config.json, adapter_config.json from the adapter repo, kept for provenance

The shards keep the base's shard names and full tensor set. Total weight size is about 19.3 GB.

Serve it with vllm.cpp

Build the vllm.cpp server (-DVLLM_CPP_SERVER=ON, target vllm-server), then:

hf download mudler/Bespoke-Nimble-9B-vllm-cpp --revision 52eead25b6723710ec16378942c9ae914d179b6f --local-dir nimble-9b
build/examples/vllm-server --model nimble-9b --served-model-name nimble --port 8000

Serve it with LocalAI

A model config for LocalAI's vllm-cpp backend:

name: nimble-9b
backend: vllm-cpp
known_usecases:
  - decisions
parameters:
  model: mudler/Bespoke-Nimble-9B-vllm-cpp
artifacts:
  - name: model
    target: model
    source:
      type: huggingface
      repo: mudler/Bespoke-Nimble-9B-vllm-cpp
      revision: 52eead25b6723710ec16378942c9ae914d179b6f

Then send the same request body to LocalAI's POST /v1/systemone with "model": "nimble-9b".

Example

Request (served on CPU from this exact directory, 2026-09-30):

curl http://localhost:8000/v1/systemone -H 'Content-Type: application/json' -d '{
  "model": "nimble",
  "state": "I was charged twice for my subscription this month. Please refund the duplicate payment.",
  "questions": {
    "refund": {"type": "noul", "instructions": "Does the user request a refund?"},
    "department": {"type": "choice", "instructions": "Which department should handle this?",
                   "criteria": {"billing": "Payments and refunds", "technical": "Software bugs", "sales": "New purchases"}},
    "urgency": {"type": "score", "instructions": "How urgent is this?",
                "criteria": ["Routine", "Urgent", "Emergency"]}}}'

Response:

{
  "model": "nimble",
  "answers": {
    "refund": {"type": "noul", "noul": 0.9966597833278789},
    "department": {"type": "choice", "choice": "billing",
                   "probabilities": {"billing": 0.9881318797907442, "technical": 0.010937236221712353, "sales": 0.0009308839875433919},
                   "confidence": 0.9383928321620487},
    "urgency": {"type": "score", "score": 0.3872836981253824,
                "legend": {"0": "Routine", "1": "Urgent", "2": "Emergency"},
                "probabilities": {"0": 0.6300001903087783, "1": 0.3527159212570611, "2": 0.017283888434160653},
                "confidence": 0.33663361186646856}
  },
  "usage": {"input_tokens": 988, "output_tokens": 0},
  "latency_ms": 75686.59
}

The answer semantics are those of Nimble's reference server (openjev): noul is p(true), choice is the most probable key, score is the expected level index, confidence is 1 - H(p) / ln(K), and probabilities are not rounded. The latency above is a 20-thread CPU and says nothing about GPU speed.

What was verified and what was not

Verified:

  • The one live request above, served by the vllm.cpp CPU server built at commit 96788348627b6a079fcc3ef7fc6676a970965d7b, returns the answers shown. They are plausible for the input.
  • Earlier on 2026-09-30, the vllm.cpp project compared this same converted directory, served on /v1/systemone on CPU, against the Bespoke authors' own parallel_schema.prepare_prompts, inference.candidate_logits and inference.decision_result over transformers 5.3.0 BF16 with the adapter unmerged through PEFT 0.21.0, also on CPU. Five requests, seven questions (choice, noul with and without criteria, score): argmax equal on 7 of 7, input token counts equal on 5 of 5, largest probability difference 0.0054. That comparison was not repeated for this upload.

Not verified:

  • No GPU (CUDA, ROCm, Metal) run. The authors' reference runner requires a CUDA BF16 GPU; the comparison above called its functions on CPU.
  • Seven questions are a check, not an accuracy measurement.
  • No vLLM gate exists, because vLLM does not serve NimbleModel decisions.
  • vllm.cpp refuses fields with more than 26 choices; the upstream release supports up to 255.
  • No GGUF or quantized variant.
Downloads last month
2
Safetensors
Model size
10B params
Tensor type
BF16
·
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for mudler/Bespoke-Nimble-9B-vllm-cpp

Finetuned
Qwen/Qwen3.5-9B
Finetuned
(971)
this model