Decision-2.0-Lux-9B-MLX-4bit (unofficial)

An unofficial 4-bit MLX conversion of vllm-sr/Decision-2.0-Lux-9B (7.94B parameters), for running on Apple Silicon. It is not made, reviewed or endorsed by the vLLM Semantic Router team or by Qwen. For the model itself — what it does, its intended use and its evaluation — see the original model card.

What was changed

Decision 2.0 is a Qwen3.5 backbone (backbone/, fully fine-tuned from Qwen/Qwen3.5-9B-Base) plus a small candidate head (decision_head.safetensors) that reads the backbone's last hidden state at the end of each option and at the end of the prompt.

  • Backbone: converted and quantized. The weights in backbone/ were renamed to the layout mlx_lm's qwen3_5 model expects and quantized to 4 bits (affine, group size 64) with mlx_lm.convert. The result is model*.safetensors + config.json in this repository. config.json declares tie_word_embeddings: true only so that no LM head is built; the LM head is never used.
  • Everything else: unchanged. decision/ is a byte-for-byte copy of the original repository without backbone/ and assets/: the candidate head, tokenizer, prompt encoding, answer normalization and the original decision2 runtime code, together with the original LICENSE and README.md.
  • New file: mlx_decision.py. Runs the backbone in MLX and everything else through the original code in decision/ (row construction, encode, collate, CandidateHead in FP32, product_answer), with temperature 1.0 as in the original package (no calibration, no score bias).

Shared prefix (on by default). Every question of a request is its own sequence that starts with the same state, as in the original. mlx_decision.py computes the longest token prefix all questions share (cut before the first option, like the original's share_context in mode="cache") once, and continues every question from a copy of that cache. This makes requests with many questions about one long state several times faster (27 questions on a ~2.5k-token state: 18.9 s → 3.8 s on an M-series Mac, same top answer on all 27, max probability difference 0.005). The arithmetic is not the exact path's, so near-tie answers can flip; set DECISION_PREFIX_CACHE=0 for the exact path (every question re-reads the state). Long requests are split into batches of at most DECISION_BATCH_TOKENS tokens (default 32768) to bound memory; this does not change answers.

Not supported: the original's mode="tree" and tau fallback, CUDA/ROCm fast kernels, and the transformers.AutoModel / pipeline entry points.

Usage

Requirements (tested): Python 3.12, mlx==0.32.3, mlx-lm==0.32.0, torch==2.14.1, transformers==5.18.0, safetensors, numpy, huggingface_hub (plus fastapi and uvicorn for the HTTP server).

hf download moritalous/Decision-2.0-Lux-9B-MLX-4bit --local-dir Decision-2.0-Lux-9B-MLX-4bit
cd Decision-2.0-Lux-9B-MLX-4bit
from mlx_decision import MLXDecision

engine = MLXDecision()  # decision/ (original files) + this repository's MLX backbone
answers, tokens = engine.system_one(
    state="The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",
    questions={
        "route": {"type": "choice", "instructions": "Which team should handle this request?",
                   "criteria": {"returns": "Refunds, replacements and damaged deliveries",
                                "billing": "Payments, invoices and charges",
                                "technical": "Product setup and faults"}},
        "receipt": {"type": "noul", "instructions": "Does the customer have a receipt?"},
    },
)
print(answers)

Or as a local System One–compatible HTTP endpoint (POST /v1/systemone):

python mlx_decision.py --port 8000

Quantization note

Measured by the converter on a 16 GB Apple Silicon Mac, not by the original authors. The official PyTorch runtime does not fit in 16 GB on a Mac, so this was compared against an MLX 8-bit conversion of the same model on a 400-question subset: top answer matched on 370 / 400 (92.5%). (For the 2B model, the MLX BF16 conversion matched the official runtime on 1,495 / 1,500.)

License

Apache-2.0, the license of the original model (LICENSE, copied unchanged from the original repository; the same text ships with Qwen/Qwen3.5-9B-Base). Original model: © the vLLM Semantic Router team; base model: © Alibaba Cloud (Qwen). The conversion, mlx_decision.py and this card are also released under Apache-2.0.

Downloads last month
79
Safetensors
Model size
8B params
Tensor type
U32
·
BF16
·
F32
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for moritalous/Decision-2.0-Lux-9B-MLX-4bit

Finetuned
Qwen/Qwen3.5-9B
Quantized
(3)
this model