DeskMind Brain 4b · 得心

得心,应手。 Brain is the decision model of DeskMind, open-source models and tools that let an agent see your screen, decide the next step and act on your own Mac. It answers System One–style typed questions (which operation next, which element, is the goal met, should I ask the user first) and returns calibrated probabilities read from answer-letter logits: nothing is generated or parsed. The server is wire-compatible with POST /v1/systemone.

This model is the strong tier of the Brain router: it answers the steps that brain-0.8b escalates (about 70% of steps in this release). It can also be served on its own.

  • Release: G18b, revision g18b-q8 (8-bit MLX, prompt format 3). main follows the current release.
  • Download: 4.5 GB.
  • Code and docs: deskmind-ai/brain. Keep deskmind.json next to the weights: it records the prompt format the model was trained with.

Revisions

Revision What it is
g18b-q8 Current release. G18b, router threshold 0.96; used by the DeskMind app v0.3.0.
g17-q8 Test build, not a release.
g14-q8 Earlier release (G14, prompt format 3, router threshold 0.94); used by the DeskMind app v0.2.0.
g13-q8 Earlier release (G13, prompt format 3).
v7b-q8 First published build (checkpoint G11b, prompt format 2).

Use

git clone https://github.com/deskmind-ai/brain && cd brain
uv sync --extra mlx
uv run hf download deskmind/brain-0.8b --revision g18b-q8 --local-dir models/brain-0.8b
uv run hf download deskmind/brain-4b --revision g18b-q8 --local-dir models/brain-4b
uv run deskmind-brain-serve --predictor mlx:models/brain-0.8b --escalate-to mlx:models/brain-4b \
  --two-stage --port 8796

The 4B can also answer every step on its own:

uv run deskmind-brain-serve --predictor mlx:models/brain-4b --port 8793 --two-stage

The 4B is 4.5 GB to download and the 0.8B 0.8 GB. If hf download fails with CAS Client Error, retry with HF_HUB_DISABLE_XET=1 in front of the command. In mainland China, ModelScope carries the same files:

uvx modelscope download --model gxcsoccer/brain-4b --revision g18b-q8 --local-dir models/brain-4b

Each reply carries a routing record: who answered, and why. Full walkthrough: deskmind-ai/brain.

Results

All numbers are our own runs; method and full tables are in docs/results.md.

Real macOS desktop, bench suite v25, 13 sandbox tasks × 3 runs, strict pass, run through the DeskMind app on an M4 Pro (48 GB):

config pass false "done" decision time p50 / p95
Router G18b (0.8B → 4B, 8-bit, threshold 0.96) 39/39 0 2.85 / 9.82 s (208 decisions)
Router G14 (earlier release, threshold 0.94) 36/39 0 0.57 / 5.25 s
  • Steps the 0.8B answers itself take p50 0.48 s / p95 0.66 s; steps escalated to the 4B take p50 3.6 s / p95 9.8 s. About 70% of steps escalate.
  • The G18b runs had the app's optional checks and notes off; there were no environment errors and no no-progress loops.
  • 39 runs over 13 tasks is a small sample, and runs cluster by task (a task tends to pass 3/3 or 0/3).
  • On the earlier suite v23, the G14 router passed 35/38 (92%) and Jev (TypeSafe, hosted) 33/38 (87%). There is no Jev run on v25.

JevBench v1.4.2, public set (231 items, the board's public_accuracy column), run locally with the official jevbench.cli:

easy (48) original (72) hard (111) public (231)
Brain 4B, G18b 48 69 76 0.835
Brain 0.8B, G18b 48 56 63 0.723
Router G18b (threshold 0.96) 48 65 71 0.797
Brain 4B, G14 48 67 85 0.866
  • Public items only. The board's headline JevBench Score also weighs 308 sealed items, calibration, speed and cost; we have not been scored on the sealed set.
  • Calibration of the G18b 4B: Brier 0.269, ECE 0.089 (G14: 0.230, 0.079).
  • Contamination check: the G18b training mix (100,703 items) shares no word 13-gram and no option set with the 231 public items (0/231).

Limitations

  • Speed: the G18b 0.8B's confidences sit in a narrow band (about 0.94–0.97), so at threshold 0.96 most steps go to the 4B and a typical decision takes about 3 s, slower than hosted models.
  • General judgement: against G14, the 4B dropped on JevBench's hard tier (85 → 76 of 111), mostly temporal/numeric, hard-judgement and multi-hop items. G18b was kept for its real-desktop reliability.
  • Saying "done": on the Chinese exact-text task the file was right in all 3 runs, but the model never said "done" and used the full 20-step budget. The grader checks the final state, so these count as passes.
  • Scope: trained and tested on macOS Finder and TextEdit sandbox tasks plus web and form decisions; untested elsewhere.
  • Probabilities are not guarantees: a probability is the model's own weighting of the options, not proof that the step is right.

Training

LoRA distillation on Qwen/Qwen3.5-4B (KL to teacher distributions plus cross-entropy to labels), merged and quantized to 8 bits. Desktop data comes from DAgger in sandboxed macOS tasks, with every visited state labelled by a desktop oracle, plus counterexamples that break label shortcuts. Web, form and evidence items come from public datasets and synthetic tasks, generated and labelled with hosted frontier-model teachers. No Jev outputs were used as labels, and the data contains no real user data. Details: docs/training.md.


中文: 得心(DeskMind)Brain 的强档:回答 0.8B 交上来的步骤(本版约占 70%),也可以单独运行。给电脑操作 agent 的每一步做带类型、带把握程度的决策,用 MLX 在 Apple Silicon 本地运行。

  • 当前发布版: G18b,版本 g18b-q8(8 位 MLX);DeskMind app v0.3.0 使用此版本。
  • 真机成绩: bench v25,13 个沙箱任务各跑 3 轮,通过 DeskMind app 运行,39/39 通过,没做完就说完成 0 次。0.8B 直接回答的步骤中位 0.48 秒;约 70% 的步骤交给 4B,中位 3.6 秒。
  • JevBench v1.4.2 公开题(231 道): 4B 0.835,0.8B 0.723,路由 0.797;训练数据与公开题无重合(0/231)。
  • 下载: hf download 加 --revision g18b-q8;国内可用 ModelScope: uvx modelscope download --model gxcsoccer/brain-4b --revision g18b-q8 --local-dir models/brain-4b。
  • 详见 deskmind-ai/brain。

License

Apache-2.0 (see LICENSE and NOTICE). Fine-tuned from Qwen/Qwen3.5-4B (Copyright Alibaba Cloud, Apache-2.0). The DeskMind name, 得心 and the logo are not covered by this licence.

Downloads last month
52
Safetensors
Model size
4B params
Tensor type
U32
·
BF16
·
MLX
Hardware compatibility
Log In to add your hardware

8-bit

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for deskmind/brain-4b

Finetuned
Qwen/Qwen3.5-4B
Quantized
(476)
this model

Collection including deskmind/brain-4b