Instructions to use olaverse/PurpleMIST-Mini-1.0 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use olaverse/PurpleMIST-Mini-1.0 with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("text-classification", model="olaverse/PurpleMIST-Mini-1.0")# pip install -U transformers accelerate # Load model directly from transformers import AutoTokenizer, AutoModel tokenizer = AutoTokenizer.from_pretrained("olaverse/PurpleMIST-Mini-1.0") model = AutoModel.from_pretrained("olaverse/PurpleMIST-Mini-1.0", device_map="auto") - Notebooks
- Google Colab
- Kaggle
PurpleMIST-Mini-1.0
PurpleMIST-Mini-1.0 Highlights
PurpleMIST-Mini-1.0 is the smallest System One decision model in the PurpleMIST family. Give it a piece of unstructured state (a message, ticket, record or transcript) and any number of typed questions about it. It returns a calibrated probability distribution for every answer, in one forward pass. It uses the same interface as PurpleMIST-Flash-1.0, at under a quarter of the size.
- 1.9B parameters. About 7.5 GB in memory, in fp32.
- 0.601 zero-shot accuracy in English on LocalLLaMA/typed-decisions (KL 0.280, Brier 0.156). That is above models up to four times its size: Bongard-mini (7.5B, 0.594) and Jeff-Gemma4-E2B (4.6B, 0.561). It is also well above Jeff-Qwen3.5-2B (2.2B, 0.511) and the input-blind Prior (0.470). It is below OpenDecider-small, a model twice its size (4B, 0.648 measured by us).
- Calibrated. ECE is 0.058 in English and 0.050 across eight African languages.
- One pass per state. All questions about a state share one sequence, and any number of options works.
- For African languages, use Flash. Mini reaches 0.416 on the translated African Typed Decisions cases, against 0.668 for Flash. A 2B backbone carries much less of these languages.
Model Overview
- Type: decision model. It scores options and does not generate text.
- Backbone: the text model of Qwen/Qwen3.5-2B-Base, with the vision tower and language-model head removed. It has 24 layers (18 Gated DeltaNet linear-attention and 6 full attention), hidden size 2048 and about 1.9B parameters.
- Head: a pointer head. Each option's logit is a scaled dot product between projections (2048 → 1024) of the answer position and the option's own line. The head has 4.2M parameters.
- Fine-tuning: full fine-tune of every weight. Flash uses LoRA instead.
- Precision: run it in fp32. Weights are stored in bf16, but the model trained with fp32 weights and bf16 matrix
multiplies. Running it entirely in bf16 costs about 3 points (0.568 in English, against 0.601).
Deciderreadsinference_dtype: float32frompurplemist_config.jsonand loads fp32 automatically. - Question types:
| Type | You give | You get back |
|---|---|---|
choice |
instructions and named options (criteria: key → description) |
{"choice": key, "probabilities": {key: p}} |
noul |
a yes/no statement | {"noul": p_yes} |
score |
instructions and ordered level descriptions (lowest first) | {"score": expected level, "probabilities": {"0": p, ...}} |
- Context: 2,048 tokens per pass. If all the questions do not fit alongside the state, they are split across several passes automatically.
Quickstart
pip install "transformers>=5.19" torch huggingface_hub
pip install flash-linear-attention causal-conv1d # optional: fast kernels for the linear-attention layers
import os, sys
from huggingface_hub import hf_hub_download
repo = "olaverse/PurpleMIST-Mini-1.0"
sys.path.insert(0, os.path.dirname(hf_hub_download(repo, "purplemist.py")))
from purplemist import Decider
d = Decider.from_pretrained(repo, device="cuda") # loads in fp32 (from the model's config)
state = {"channel": "email",
"message": "Hi, I was charged twice for the same order last night. Please refund the extra payment."}
questions = {
"team": {"type": "choice", "instructions": "Which team should handle this?",
"criteria": {"billing": "payments, charges and refunds", "delivery": "orders in transit",
"technical": "app or account problems"}},
"refund": {"type": "noul", "instructions": "The customer is asking for money back."},
"urgency": {"type": "score", "instructions": "How urgent is this?",
"criteria": ["can wait", "normal queue", "today", "immediately"]},
}
print(d.decide(state, questions))
# {"team": {"choice": ..., "probabilities": {...}}, "refund": {"noul": ...}, "urgency": {"score": ..., "probabilities": {...}}}
state can be a string or any JSON-serialisable object.
Serving over HTTP
serve.py runs the model as a small local server with the System One request shape (POST /v1/systemone), so
clients built for that API can call it. It needs only the packages from the quickstart.
hf download olaverse/PurpleMIST-Mini-1.0 serve.py purplemist.py --local-dir purplemist
python purplemist/serve.py --model olaverse/PurpleMIST-Mini-1.0 --port 8000
curl -s localhost:8000/v1/systemone -H 'content-type: application/json' -d '{
"state": {"message": "I was charged twice, please refund me"},
"questions": {
"team": {"type": "choice", "instructions": "Which team?", "criteria": {"billing": "payments", "tech": "app problems"}},
"refund": {"type": "noul", "instructions": "The customer asks for money back."}}}'
# {"model": "...", "answers": {"team": {"choice": ..., "probabilities": {...}}, "refund": {"noul": ...}}, "latency_ms": ...}
The server also offers GET /health and GET /v1/models. It handles one request at a time. --max-len raises the
2,048-token limit per pass, for very long states or questions with hundreds of options. The model was trained on
2,048 tokens, so answers on much longer inputs may be weaker. Ollama and llama.cpp cannot run this model: they
serve text generators, and PurpleMIST scores options with its own head instead of generating text.
Evaluation
All models were run through one harness. Each request contains a whole case: the state and all of its questions. Every answer is turned into a probability vector over the same options, and every model is scored by the same code. LLMs were prompted for JSON probabilities at temperature 0 with reasoning off.
LocalLLaMA/typed-decisions (English, zero-shot)
These are the zero-shot rows from the benchmark card, with our measured rows added. The test split has 400 cases and 2,000 decisions.
| Model | Size | Accuracy ↑ | KL ↓ | Brier ↓ | ECE ↓ | Source |
|---|---|---|---|---|---|---|
| meraGPT Decider 1 | undisclosed | 0.768 | 0.096 | 0.052 | 0.180 | benchmark card |
| Liquid AI d1 | undisclosed | 0.742 | 0.475 | 0.155 | 0.124 | benchmark card |
| TypeSafe Jev 1.13.0 | undisclosed | 0.727 | 1.442 | 0.148 | 0.144 | benchmark card |
| PurpleMIST-Flash-1.0 | 7.9B | 0.720 | 0.189 | 0.098 | 0.056 | measured by us |
| DeepSeek-V4-Pro (prompted) | 1.6T | 0.720 | 1.359 | 0.185 | 0.108 | measured by us |
| Featherless Simple Jev | 35B (3B active) | 0.716 | 0.488 | 0.176 | – | benchmark card |
| Qwen3.8-Flash (prompted) | undisclosed | 0.712 | 0.781 | 0.172 | 0.078 | measured by us |
| prima-ratio + 12B | 12B | 0.702 | 0.564 | 0.234 | 0.146 | benchmark card (self-reported) |
| OpenDecider-small | 4B | 0.671 | 0.211 | 0.117 | – | benchmark card (self-reported) |
| OpenDecider-small | 4B | 0.648 | 0.224 | 0.126 | 0.059 | measured by us |
| PurpleMIST-Mini-1.0 | 1.9B | 0.601 | 0.280 | 0.156 | 0.058 | measured by us |
| Bongard-mini | 7.5B | 0.594 | 0.256 | 0.132 | 0.067 | benchmark card (self-reported) |
| Jeff-Gemma4-E2B | 4.6B | 0.561 | 0.403 | 0.219 | 0.188 | benchmark card |
| Jeff-Qwen3.5-2B | 2.2B | 0.511 | 0.460 | 0.237 | 0.203 | benchmark card |
| Jeff-Qwen3.5-0.8B | 0.85B | 0.483 | 0.679 | 0.313 | 0.251 | benchmark card |
| Prior (ignores the input) | – | 0.470 | 0.347 | 0.189 | 0.088 | benchmark card |
Sizes are total parameters; "undisclosed" means the provider does not publish one. Submitters compute ECE differently, so compare KL and Brier rather than ECE.
African Typed Decisions
| Model | Size | English original | African, translated | English → African drop | African, native | KL ↓ | Brier ↓ | ECE ↓ |
|---|---|---|---|---|---|---|---|---|
| DeepSeek-V4-Pro (prompted) | 1.6T | 0.720 | 0.669 | 0.051 | 0.955 † | 1.125 | 0.199 | 0.121 |
| PurpleMIST-Flash-1.0 | 7.9B | 0.720 | 0.668 | 0.053 | 0.932 ‡ | 0.214 | 0.115 | 0.049 |
| Qwen3.8-Flash (prompted) | undisclosed | 0.712 | 0.643 | 0.070 | 0.916 † | 0.862 | 0.216 | 0.132 |
| OpenDecider-small | 4B | 0.648 | 0.477 | 0.171 | 0.635 | 0.365 | 0.199 | 0.027 |
| PurpleMIST-Mini-1.0 | 1.9B | 0.601 | 0.416 | 0.185 | 0.769 ‡ | 0.399 | 0.221 | 0.050 |
| laya-multilingual | 0.32B | 0.348 | 0.322 | 0.026 | 0.280 | 5.053 | 0.545 | 0.360 |
KL, Brier and ECE are on the translated set: 3,010 cases, 15,050 decisions.
† DeepSeek-V4-Pro wrote the Nigerian-language native cases, and Qwen3.8-Flash wrote or checked all of them, so both are scored partly against their own answers on that column. ‡ PurpleMIST was trained on native cases generated by the same pipeline. They were separate cases from those in the benchmark.
By language (accuracy, translated cases):
| Model | Amharic | Hausa | Igbo | Pidgin | Somali | Swahili | Yorùbá | isiZulu |
|---|---|---|---|---|---|---|---|---|
| PurpleMIST-Flash-1.0 | 0.649 | 0.649 | 0.654 | 0.710 | 0.683 | 0.690 | 0.644 | 0.671 |
| OpenDecider-small | 0.491 | 0.421 | 0.473 | 0.619 | 0.443 | 0.510 | 0.439 | 0.456 |
| PurpleMIST-Mini-1.0 | 0.330 | 0.414 | 0.422 | 0.589 | 0.371 | 0.418 | 0.420 | 0.406 |
| laya-multilingual | 0.334 | 0.326 | 0.317 | 0.351 | 0.299 | 0.311 | 0.352 | 0.299 |
Mini does best on Nigerian Pidgin, which is closest to English, and worst on Amharic, the one language here not written in Latin script.
Training
Data: the same training data as Flash. That is 300,000 states sampled from a 1.71M-decision set, with every African state included. The sources are:
- public System One and preference data: tasksource, Open-Jev
train, HelpSteer2, jev-decisions-clean50k, two synthetic sets and Najd development; - 86,134 African decisions, written natively in or translated into eight languages.
Every source allows commercial use. LocalLLaMA/typed-decisions
trainwas not used, which keeps the English result zero-shot.- public System One and preference data: tasksource, Open-Jev
Objective: soft-target cross-entropy against each question's reference distribution. Options are shuffled during training.
Run: full fine-tune for 6,007 steps, with fp32 master weights and bf16 autocast. The learning rate was 2e-5 and weight decay 0.01. Sequences were up to 2,048 tokens.
Calibration: one temperature per question type, fitted in fp32 on 2,751 states (5,345 questions) that training never used. The states are spread across eight sources, each weighted equally. The temperatures are large: choice 24.5, noul 33.5, score 23.5. Mini's raw scores are very sharp, and the temperatures bring them back to calibrated probabilities.
Deciderapplies them for you. If you read raw logits yourself, divide by these.
Limitations
- African languages: 0.33–0.59 on translated cases. Use PurpleMIST-Flash-1.0 for these languages.
- Gold comes from LLM teachers. Reference answers were generated by language models, so a score measures agreement with those teachers, not ground truth.
- Text only. Images in the state are not read.
- bf16: running in bf16 lowers accuracy. Keep the default fp32.
eval/files: these come from the training run's evaluation, in bf16 and before recalibration. The numbers on this card replace them.- Score questions:
scorereturns the expected level. Useprobabilitiesif you need the most likely level.
Citation
@misc{purplemist-mini-1.0,
title = {PurpleMIST-Mini-1.0: a small calibrated System One decision model},
author = {Olaverse},
year = {2026},
url = {https://huggingface.co/olaverse/PurpleMIST-Mini-1.0}
}
- Downloads last month
- -
Model tree for olaverse/PurpleMIST-Mini-1.0
Base model
Qwen/Qwen3.5-2B-BaseDatasets used to train olaverse/PurpleMIST-Mini-1.0
olaverse/african-typed-decisions
Collection including olaverse/PurpleMIST-Mini-1.0
Evaluation results
- LocalLLaMA/typed-decisions leaderboard
- Accuracy View evaluation resultssource
General, zero-shot: never trained on the Typed Decisions train split. Test split, 400 cases, 2,000 decisions, one request per case with the state and all five questions, all answered, zero errors. Accuracy = agreement with the gold label; KL = KL(gold || prediction); Brier summed over options. One forward pass per state, no generated tokens. fp32, calibrated temperatures from purplemist_config.json.0.6 * - Kl From Gold View evaluation resultssource
General, zero-shot: never trained on the Typed Decisions train split. Test split, 400 cases, 2,000 decisions, one request per case with the state and all five questions, all answered, zero errors. Accuracy = agreement with the gold label; KL = KL(gold || prediction); Brier summed over options. One forward pass per state, no generated tokens. fp32, calibrated temperatures from purplemist_config.json.0.28 * - Brier View evaluation resultssource
General, zero-shot: never trained on the Typed Decisions train split. Test split, 400 cases, 2,000 decisions, one request per case with the state and all five questions, all answered, zero errors. Accuracy = agreement with the gold label; KL = KL(gold || prediction); Brier summed over options. One forward pass per state, no generated tokens. fp32, calibrated temperatures from purplemist_config.json.0.16 * - .eval_results/african-typed-decisions.yaml error

