Instructions to use moritalous/Decision-2.0-Lux-9B-MLX-4bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use moritalous/Decision-2.0-Lux-9B-MLX-4bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download moritalous/Decision-2.0-Lux-9B-MLX-4bit --local-dir Decision-2.0-Lux-9B-MLX-4bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
Decision-2.0-Lux-9B-MLX-4bit (unofficial)
An unofficial 4-bit MLX conversion of vllm-sr/Decision-2.0-Lux-9B (7.94B parameters), for running on Apple Silicon. It is not made, reviewed or endorsed by the vLLM Semantic Router team or by Qwen. For the model itself — what it does, its intended use and its evaluation — see the original model card.
What was changed
Decision 2.0 is a Qwen3.5 backbone (backbone/, fully fine-tuned from Qwen/Qwen3.5-9B-Base)
plus a small candidate head (decision_head.safetensors) that reads the backbone's last hidden state at the
end of each option and at the end of the prompt.
- Backbone: converted and quantized. The weights in
backbone/were renamed to the layoutmlx_lm'sqwen3_5model expects and quantized to 4 bits (affine, group size 64) withmlx_lm.convert. The result ismodel*.safetensors+config.jsonin this repository.config.jsondeclarestie_word_embeddings: trueonly so that no LM head is built; the LM head is never used. - Everything else: unchanged.
decision/is a byte-for-byte copy of the original repository withoutbackbone/andassets/: the candidate head, tokenizer, prompt encoding, answer normalization and the originaldecision2runtime code, together with the originalLICENSEandREADME.md. - New file:
mlx_decision.py. Runs the backbone in MLX and everything else through the original code indecision/(row construction,encode,collate,CandidateHeadin FP32,product_answer), with temperature 1.0 as in the original package (no calibration, no score bias).
Shared prefix (on by default). Every question of a request is its own sequence that starts with the
same state, as in the original. mlx_decision.py computes the longest token prefix all questions share
(cut before the first option, like the original's share_context in mode="cache") once, and continues
every question from a copy of that cache. This makes requests with many questions about one long state
several times faster (27 questions on a ~2.5k-token state: 18.9 s → 3.8 s on an M-series Mac, same top answer on
all 27, max probability difference 0.005). The arithmetic is not the exact path's, so near-tie answers can flip;
set DECISION_PREFIX_CACHE=0 for the exact path (every question re-reads the state). Long requests are split into
batches of at most DECISION_BATCH_TOKENS tokens (default 32768) to bound memory; this does not change answers.
Not supported: the original's mode="tree" and tau fallback, CUDA/ROCm fast kernels, and the
transformers.AutoModel / pipeline entry points.
Usage
Requirements (tested): Python 3.12, mlx==0.32.3, mlx-lm==0.32.0, torch==2.14.1, transformers==5.18.0,
safetensors, numpy, huggingface_hub (plus fastapi and uvicorn for the HTTP server).
hf download moritalous/Decision-2.0-Lux-9B-MLX-4bit --local-dir Decision-2.0-Lux-9B-MLX-4bit
cd Decision-2.0-Lux-9B-MLX-4bit
from mlx_decision import MLXDecision
engine = MLXDecision() # decision/ (original files) + this repository's MLX backbone
answers, tokens = engine.system_one(
state="The order arrived damaged yesterday. The customer has a receipt and asks for a replacement today.",
questions={
"route": {"type": "choice", "instructions": "Which team should handle this request?",
"criteria": {"returns": "Refunds, replacements and damaged deliveries",
"billing": "Payments, invoices and charges",
"technical": "Product setup and faults"}},
"receipt": {"type": "noul", "instructions": "Does the customer have a receipt?"},
},
)
print(answers)
Or as a local System One–compatible HTTP endpoint (POST /v1/systemone):
python mlx_decision.py --port 8000
Quantization note
Measured by the converter on a 16 GB Apple Silicon Mac, not by the original authors. The official PyTorch runtime does not fit in 16 GB on a Mac, so this was compared against an MLX 8-bit conversion of the same model on a 400-question subset: top answer matched on 370 / 400 (92.5%). (For the 2B model, the MLX BF16 conversion matched the official runtime on 1,495 / 1,500.)
License
Apache-2.0, the license of the original model (LICENSE, copied unchanged from the original repository;
the same text ships with Qwen/Qwen3.5-9B-Base). Original model: © the vLLM Semantic Router team; base model: © Alibaba Cloud
(Qwen). The conversion, mlx_decision.py and this card are also released under Apache-2.0.
- Downloads last month
- 79
4-bit
Model tree for moritalous/Decision-2.0-Lux-9B-MLX-4bit
Base model
Qwen/Qwen3.5-9B-Base