Instructions to use manjunathshiva/opendecider-small-mlx-8bit with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use manjunathshiva/opendecider-small-mlx-8bit with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] huggingface-cli download --local-dir opendecider-small-mlx-8bit manjunathshiva/opendecider-small-mlx-8bit
- Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Atomic Chat
OpenDecider-small MLX 8-bit
OpenDecider-small, the 4B decision model, merged and
quantised to 8-bit for Apple Silicon with MLX: 4.0 GB instead of 8 GB, and 4.5 GB of
memory while answering. Ask typed questions (choice, score, noul) about any text or JSON and get a
calibrated probability for every option. Apache-2.0.
Recommended Mac build: same answers as the full model on 399 of 400 general and 1,955 of 2,000 typed-decisions questions, at half the memory and about 2× the speed of PyTorch on Apple Silicon.
Installation
pip install "opendecider[mlx]"
Quickstart
from opendecider import load
model = load("manjunathshiva/opendecider-small-mlx-8bit")
r = model.system_one(
"Hi, we were billed twice for March. Please refund the duplicate today or we will cancel our plan.",
{"department": {"type": "choice", "instructions": "Which department should handle this?",
"criteria": {"billing": "invoices, payments, refunds", "technical": "bugs, outages", "other": "everything else"}},
"churn_risk": {"type": "noul", "instructions": "Does the user threaten to cancel or leave?"}})
print(r["answers"]["department"]["choice"], r["answers"]["churn_risk"]["noul"])
Same model, smaller: how close is it?
Scored with the benchmark harness on the same questions as the full-precision release:
| benchmark | OpenDecider-small (bf16, PyTorch) | MLX 8-bit | same top answer as bf16 |
|---|---|---|---|
| 200 general decisions | 0.735 | 0.730 | 399/400 |
| typed-decisions (2,000 decisions) | 0.672 | 0.673 | 1955/2000 |
Will it fit?
| Mac | memory used | latency, one question |
|---|---|---|
| Apple M4 Max, 64 GB | 4.5 GB | 66 ms (short question); 148 ms median on benchmark questions |
Any Apple Silicon Mac with 8 GB or more should run it (not every size tested).
Links
- Full-precision model and benchmarks: https://huggingface.co/manjunathshiva/opendecider-small
- GitHub: https://github.com/manjunathshiva/opendecider
- Collection: https://huggingface.co/collections/manjunathshiva/opendecider-6ab8c838909092518d50a9ea
Apache 2.0 · Base model Qwen3-4B-Instruct-2507 (Apache-2.0) · Manjunath Janardhan
- Downloads last month
- 15
8-bit
Model tree for manjunathshiva/opendecider-small-mlx-8bit
Base model
Qwen/Qwen3-4B-Instruct-2507