Instructions to use Orbalapp/orb-1-mini with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- MLX
How to use Orbalapp/orb-1-mini with MLX:
# Make sure mlx-lm is installed # pip install --upgrade mlx-lm # Generate text with mlx-lm from mlx_lm import load, generate model, tokenizer = load("Orbalapp/orb-1-mini") prompt = "Write a story about Einstein" messages = [{"role": "user", "content": prompt}] prompt = tokenizer.apply_chat_template( messages, add_generation_prompt=True ) text = generate(model, tokenizer, prompt=prompt, verbose=True) - Notebooks
- Google Colab
- Kaggle
- Local Apps Settings
- LM Studio
- Pi
How to use Orbalapp/orb-1-mini with Pi:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Orbalapp/orb-1-mini"
Configure the model in Pi
# Install Pi: npm install -g @earendil-works/pi-coding-agent # Add to ~/.pi/agent/models.json: { "providers": { "mlx-lm": { "baseUrl": "http://localhost:8080/v1", "api": "openai-completions", "apiKey": "none", "models": [ { "id": "Orbalapp/orb-1-mini" } ] } } }Run Pi
# Start Pi in your project directory: pi
- MLX LM
How to use Orbalapp/orb-1-mini with MLX LM:
Generate or start a chat session
# Install MLX LM uv tool install mlx-lm # Interactive chat REPL mlx_lm.chat --model "Orbalapp/orb-1-mini"
Run an OpenAI-compatible server
# Install MLX LM uv tool install mlx-lm # Start the server mlx_lm.server --model "Orbalapp/orb-1-mini" # Calling the OpenAI-compatible server with curl curl -X POST "http://localhost:8000/v1/chat/completions" \ -H "Content-Type: application/json" \ --data '{ "model": "Orbalapp/orb-1-mini", "messages": [ {"role": "user", "content": "Hello"} ] }' - Hermes Agent
How to use Orbalapp/orb-1-mini with Hermes Agent:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Orbalapp/orb-1-mini"
Configure Hermes
# Install Hermes: curl -fsSL https://hermes-agent.nousresearch.com/install.sh | bash hermes setup # Point Hermes at the local server: hermes config set model.provider custom hermes config set model.base_url http://127.0.0.1:8080/v1 hermes config set model.default Orbalapp/orb-1-mini
Run Hermes
hermes
- Atomic Chat
- OpenClaw
How to use Orbalapp/orb-1-mini with OpenClaw:
Start the MLX server
# Install MLX LM: uv tool install mlx-lm # Start a local OpenAI-compatible server: mlx_lm.server --model "Orbalapp/orb-1-mini"
Configure OpenClaw
# Install OpenClaw: npm install -g openclaw@latest # Register the local server and set it as the default model: openclaw onboard --non-interactive --mode local \ --auth-choice custom-api-key \ --custom-base-url http://127.0.0.1:8080/v1 \ --custom-model-id "Orbalapp/orb-1-mini" \ --custom-provider-id mlx-lm \ --custom-compatibility openai \ --custom-text-input \ --accept-risk \ --skip-health
Run OpenClaw
openclaw agent --local --agent main --message "Hello from Hugging Face"
orb-1-mini
orb-1-mini cleans up dictation. It takes the text a speech engine wrote down and returns the same message, cleaned up: fillers, repeated words and false starts removed, self-corrections resolved, punctuation, capitals and numbers fixed. It keeps the speaker's words and their order. It does not shorten, summarise, answer questions or add anything.
It is the on-device model in Orbal, a dictation app for the Mac. A 4-bit MLX fine-tune of Qwen3.5-0.8B, 0.42 GB, for Apple silicon.
| Dictated | orb-1-mini |
|---|---|
| So um the meeting is on Tuesday, no, Thursday at, uh, three thirty, and it's like twenty five percent over budget. | The meeting is on Thursday at 3:30, and it's 25% over budget. |
| So um can you, can you move the the logo to, sorry, to this side? | Can you move the logo to this side? |
| It costs twelve ninety nine a month, so that's about fifty dollars for four months. | It costs $12.99 a month, so that's about $50 for 4 months. |
Versions
v2(round 5, current): writes numbers, times, dates, money and percentages as people type them.v1(round 4): the first release.
Pin a version with revision="v2".
Prompts
The model was trained on these system prompts. Use them exactly as written.
| Use | System prompt |
|---|---|
| Clean up | Shape this dictation. |
| For a chat message | Shape this dictation for a chat message. |
| For an email | Shape this dictation for an email. |
| For an AI agent | Shape this dictation for an AI agent. |
| For a document | Shape this dictation for a document. |
When names must be spelled a certain way, add a line to the system prompt:
Spell these names exactly as written: Anna, Orbal.
The user message is the dictated text alone. Turn thinking off, decode greedily, and cap the answer at the input's length in tokens plus 16.
Use with mlx-lm
from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler
model, tokenizer = load("Orbalapp/orb-1-mini")
text = "So um can you, can you move the the logo to, sorry, to this side?"
messages = [{"role": "system", "content": "Shape this dictation."},
{"role": "user", "content": text}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True,
enable_thinking=False, tokenize=False)
cap = len(tokenizer.encode(text)) + 16
print(generate(model, tokenizer, prompt, max_tokens=cap, sampler=make_sampler(temp=0.0)))
# Can you move the logo to this side?
Results
Tested on 178 English dictations, 171 real ones from Orbal's developer and 7 made up, and on 15 made-up dictations with numbers said as words. None were used in training.
| v2, clean up | v2, the four other prompts | v1, clean up | |
|---|---|---|---|
| Passes Orbal's check that nothing was added | 100% | 99% to 100% | 100% |
| Meaning kept, blind check by Qwen 3.8 27B | 99% | 93% to 98% | 100% |
| Content words kept | 99% | 98% to 99% | 100% |
| Answers with a word added (of 178) | 0 | 0 to 4 | 0 |
| Names made up | 0 | 0 | 0 |
| Questions answered | 0 | 0 | 0 |
| Numbers typed as digits (22 in the numbers test) | 22 | 19 |
Speed with mlx-lm on an M1 Pro: it loads in 1.2 s; a typical dictation takes 0.13 s, and 9 in 10 take under 0.47 s. It keeps every word, so the time grows with the length: a 347-word dictation took 5.7 s.
What it does badly
- Numbers over 99 and forms like "half past two" were left out of training, because Orbal's check cannot yet verify their digits. The model may leave them as said.
- The four style prompts change very little. Email sometimes adds paragraph breaks; otherwise they stay close to the clean-up prompt.
- Long dictations are slow, see above.
- English only. Other languages were not tested.
How it was trained
Distilled from Qwen 3.8 27B. The teacher wrote 8,323 example messages and
a spoken version of each, with fillers, false starts, self-corrections and,
for v2, numbers said as words.
A third were read aloud with macOS say and transcribed by Apple's speech
engine, for real mishearings, and all went through Orbal's own text rules.
The teacher then cleaned each one up five times, once per prompt. The
24,920 answers that passed every check (nothing added, every name, number,
negation and pointing word kept) trained a LoRA, rank 16, for 2 epochs. It
was merged and quantised to 4-bit MLX. No user dictations were used for
training.
License and credit
Apache-2.0, like Qwen3.5-0.8B, which it is built on (Copyright 2026 Alibaba Cloud). Orbal changed it: fine-tuned, merged and quantised. If you share orb-1-mini or a model made from it, keep the NOTICE file and credit "orb-1-mini by Orbal".
- Downloads last month
- -
4-bit