orb-1-mini

orb-1-mini cleans up dictation. It takes the text a speech engine wrote down and returns the same message, cleaned up: fillers, repeated words and false starts removed, self-corrections resolved, punctuation, capitals and numbers fixed. It keeps the speaker's words and their order. It does not shorten, summarise, answer questions or add anything.

It is the on-device model in Orbal, a dictation app for the Mac. A 4-bit MLX fine-tune of Qwen3.5-0.8B, 0.42 GB, for Apple silicon.

Dictated orb-1-mini
So um the meeting is on Tuesday, no, Thursday at, uh, three thirty, and it's like twenty five percent over budget. The meeting is on Thursday at 3:30, and it's 25% over budget.
So um can you, can you move the the logo to, sorry, to this side? Can you move the logo to this side?
It costs twelve ninety nine a month, so that's about fifty dollars for four months. It costs $12.99 a month, so that's about $50 for 4 months.

Versions

  • v2 (round 5, current): writes numbers, times, dates, money and percentages as people type them.
  • v1 (round 4): the first release.

Pin a version with revision="v2".

Prompts

The model was trained on these system prompts. Use them exactly as written.

Use System prompt
Clean up Shape this dictation.
For a chat message Shape this dictation for a chat message.
For an email Shape this dictation for an email.
For an AI agent Shape this dictation for an AI agent.
For a document Shape this dictation for a document.

When names must be spelled a certain way, add a line to the system prompt: Spell these names exactly as written: Anna, Orbal.

The user message is the dictated text alone. Turn thinking off, decode greedily, and cap the answer at the input's length in tokens plus 16.

Use with mlx-lm

from mlx_lm import load, generate
from mlx_lm.sample_utils import make_sampler

model, tokenizer = load("Orbalapp/orb-1-mini")
text = "So um can you, can you move the the logo to, sorry, to this side?"
messages = [{"role": "system", "content": "Shape this dictation."},
            {"role": "user", "content": text}]
prompt = tokenizer.apply_chat_template(messages, add_generation_prompt=True,
                                       enable_thinking=False, tokenize=False)
cap = len(tokenizer.encode(text)) + 16
print(generate(model, tokenizer, prompt, max_tokens=cap, sampler=make_sampler(temp=0.0)))
# Can you move the logo to this side?

Results

Tested on 178 English dictations, 171 real ones from Orbal's developer and 7 made up, and on 15 made-up dictations with numbers said as words. None were used in training.

v2, clean up v2, the four other prompts v1, clean up
Passes Orbal's check that nothing was added 100% 99% to 100% 100%
Meaning kept, blind check by Qwen 3.8 27B 99% 93% to 98% 100%
Content words kept 99% 98% to 99% 100%
Answers with a word added (of 178) 0 0 to 4 0
Names made up 0 0 0
Questions answered 0 0 0
Numbers typed as digits (22 in the numbers test) 22 19

Speed with mlx-lm on an M1 Pro: it loads in 1.2 s; a typical dictation takes 0.13 s, and 9 in 10 take under 0.47 s. It keeps every word, so the time grows with the length: a 347-word dictation took 5.7 s.

What it does badly

  • Numbers over 99 and forms like "half past two" were left out of training, because Orbal's check cannot yet verify their digits. The model may leave them as said.
  • The four style prompts change very little. Email sometimes adds paragraph breaks; otherwise they stay close to the clean-up prompt.
  • Long dictations are slow, see above.
  • English only. Other languages were not tested.

How it was trained

Distilled from Qwen 3.8 27B. The teacher wrote 8,323 example messages and a spoken version of each, with fillers, false starts, self-corrections and, for v2, numbers said as words. A third were read aloud with macOS say and transcribed by Apple's speech engine, for real mishearings, and all went through Orbal's own text rules. The teacher then cleaned each one up five times, once per prompt. The 24,920 answers that passed every check (nothing added, every name, number, negation and pointing word kept) trained a LoRA, rank 16, for 2 epochs. It was merged and quantised to 4-bit MLX. No user dictations were used for training.

License and credit

Apache-2.0, like Qwen3.5-0.8B, which it is built on (Copyright 2026 Alibaba Cloud). Orbal changed it: fine-tuned, merged and quantised. If you share orb-1-mini or a model made from it, keep the NOTICE file and credit "orb-1-mini by Orbal".

Downloads last month
-
Safetensors
Model size
0.8B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

4-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Orbalapp/orb-1-mini

Finetuned
(435)
this model