- Muse-Glimmer-30B (text decoder) β Apple Core AI (
.aimodel)
Muse-Glimmer-30B (text decoder) β Apple Core AI (.aimodel)
Meta's Muse-Glimmer-30B converted to Apple's Core AI (the Core ML successor announced at WWDC26), ready to run on macOS 27. This is the text decoder of their 30B agentic VLM β the perception encoder is not included.
It decodes faster than Meta's own on-device Apple-GPU build of the same model.
| weights | decode tok/s | prompt tok/s | |
|---|---|---|---|
ExecuTorch k-quant-17G-128K-text-solo-metal (Meta's published figure) |
17.9 GB | 23.7 | β |
this bundle β Core AI int4hu |
16.35 GB | 26.69 | 269.0 |
+12.6% decode on 8.7% fewer bytes, on Apple's stock coreai-pipelined GPU engine with no
custom Metal kernels.
Mac Studio M4 Max (40-core GPU, 128 GB, macOS 27.0 26A5406e), 512 prompt / 1024 generation /
3 trials, llm-benchmark. Meta's figure is theirs, not a re-measurement: batch 1, greedy,
averaged over a prompt set they do not publish. Their M4 Max is the same 546 GB/s bin β
23.7 Γ 17.9 GB is 424 GB/s of traffic, which the 410 GB/s bin cannot produce. Chassis differs and is worth stating in a thermal-aware comparison: theirs is a MacBook, this is a Mac Studio β same chip, better sustained cooling here. The same-machine table below is unaffected (every arm ran on this Mac Studio); only the row against their published figure carries that asymmetry. Their metal
backend is MLX-native, per their own README.
Decode barely moves with context β 27.46 tok/s at 128 prompt tokens, 26.73 at 2048 β which is what 39 of 52 layers capped at a 2048-token window should look like.
Both artifacts run on the same machine, same prompts, greedy, batch 1, 192 new tokens, interleaved A/B/A/B with a 45 s cooldown between every run (block-ordering is not safe here β ExecuTorch decayed 23.5 β 17.4 tok/s inside a single block, so whoever runs second inherits a hot GPU):
| prompt | ExecuTorch | Core AI |
|---|---|---|
| 1 (code + explanation) | 19.3 / 19.7 | 16.2 (cold) / 27.5 |
| 2 (step-by-step reasoning) | 24.0 / 23.9 | 27.3 / 27.4 |
| 3 (long-context tradeoff) | 1 token, stopped | 27.7 / 27.7 |
Core AI holds 27.3β27.7 across every prompt and round. ExecuTorch ranges 19.3β24.0; at its own
best prompt it reproduces Meta's published 23.7 almost exactly, and Core AI is +14% there.
Prompt 1 round 1 for Core AI is a cold-cache artifact (ExecuTorch had just filled the page
cache with 17.9 GB) and is discarded on the strength of round 2; on prompt 3 the ExecuTorch
runner emitted one token despite --ignore_eos=true, which is a runner behaviour and is not
counted as a win.
Speculative decoding on this bundle: 1.3β2.0Γ, lossless, no drafter
Meta's DFlash figure (37.8 tok/s) buys speculation with a separate 5.1 GB block-diffusion
drafter β 31% more bytes. The same lever works on this bundle for zero extra bytes: the
decode graph runs its head on every position and takes a dynamic input_ids, so K drafted
tokens verify in one forward, and an n-gram (prompt-lookup) drafter needs no weights at all.
No re-export; the bundle you download is the one these numbers were measured on.
256 generated tokens, greedy, batch 1, best draft length per workload:
| workload | spec off | spec on | vs DFlash 37.8 | |
|---|---|---|---|---|
| free chat | 27.31 | 36.65 | 1.34Γ | 0.97Γ |
| code rewrite | 27.37 | 53.71 | 1.96Γ | 1.42Γ |
| tool calling (ATEM, 3 tools) | 27.25 | 50.19 | 1.84Γ | 1.33Γ |
Every committed token is the model's own greedy argmax, so output is unchanged β 46 A/B runs, 46/46 byte-identical to the same loop with drafting off.
Stated plainly, because n-gram drafting is workload-bound: these three prompts are not a distribution, and Meta's 37.8 is an average over a prompt set they do not publish, so treat the comparison as directional. The advantage also decays with generation length β at 512 tokens code falls to 1.40Γ and free chat to 30.10 tok/s, below their number; tool calling holds (1.85Γ) because the ATEM protocol keeps quoting the prompt for the whole turn. Draft length matters more than acceptance does: verify cost here is a staircase (S β€ 3 free, S = 4β¦8 ~1.47Γ, S β₯ 9 ~2.3Γ), so K=8 makes free chat 5% slower while K=2 makes it 34% faster.
Raw MLX is the third arm, and it changes what the ExecuTorch result means
ExecuTorch's metal backend is MLX-native, so beating it could mean beating MLX or beating
the wrapper around MLX. Only raw MLX separates the two. Same machine, same prompts, greedy,
192 tokens, interleaved CA/ET/MLX with a 45 s cooldown between every run:
| p1 r1 | p1 r2 | p2 r1 | p2 r2 | mean | |
|---|---|---|---|---|---|
Core AI int4hu, 16.35 GB |
27.5 | 27.1 | 27.5 | 27.6 | 27.43 |
MLX mlx-community/β¦-4bit, 18 GB |
27.18 | 27.38 | 27.50 | 27.37 | 27.36 |
ExecuTorch k-quant-17Gβ¦metal, 17.9 GB |
24.0 | 23.8 | 24.1 | 24.1 | 24.00 |
Core AI and raw MLX are indistinguishable (+0.3%). Both beat Meta's own build by ~14%.
So the honest reading is not "Core AI is fast here" β it is that Meta's shipped on-device
artifact leaves ~14% on the table against the runtime it is built on. Core AI matching MLX
at 27.9 B dense is what coreai-vs-mlx-speed.md
already predicts: Core AI β₯ MLX on small dense, converging to a tie as the model grows and
MLX's 4-bit byte advantage cashes in. This is the largest dense point on that curve so far,
and it lands on the tie.
Prompt processing is not matched here and no claim is made from it (Core AI 269 tok/s at 512 prompt tokens, MLX 128β136 at 77β81 β different lengths, different batching).
Not claimed here: Meta's published quality figure (1.0% degradation across 15 benchmarks for the 17G quant) has no matched counterpart; this port has a token-exact gate and read generations, not a benchmark suite.
The architecture, and what the port had to do about it
52 layers, hidden 6656, GQA 32 q / 2 kv heads (head_dim 128), SwiGLU intermediate 19968, vocab 202 048 with an untied lm_head, 131 072 context β 27.855 B parameters in the text tower alone. Gemma-3 shaped, with four things that are not:
sliding(2048) Γ 3 + full Γ 1layer pattern where the full layers are NoPE βlayer_rope_theta[i] == 0marks them, and only the sliding layers carry rotary.- A sigmoid output gate on every attention layer, reading the same pre-attention hidden
states as Q/K/V β so the re-authored module folds all four into one
qkvg_proj(4 GEMMs β 1). - Weight-less RMSNorm on Q/K and on the embedding output.
- Two epsilons across the sandwich norms (1e-5 pre, 1e-8 post), and logits pre-scaled by
output_multiplierbefore a tanh softcap at 20.
qk_scale_factor (3.87) is folded into the SDPA scale rather than applied to Q β attention is
softmax((aQ)Β·Kα΅/βd), a is a scalar and rotary is a rotation, so it is algebraically
identical and saves a full-width multiply per layer.
The bundle needed a stop token declared
Shipped and then fixed: the first published bundle never terminated. The model
answers correctly, emits <|eot|>, and then loops to=self/to=user re-emitting
the same answer until the token budget runs out β 2048 tokens where 491 were needed.
The runtime resolves extra stop tokens from tokenizer_config.json
(additional_special_tokens, an array-valued eos_token, or added_tokens_decoder
entries matching end_of_turn / im_end / eot_id / endoftext / eot_token).
This checkpoint offers none of them: its eos_token is the plain string
<|end_of_text|> and its turn ends with <|eot|> (200008), which matches no
pattern in that list. The fact is upstream β generation_config.json declares
eos_token_id: [200001, 200008] β but that file is not part of a bundle and the
runtime does not read it.
The exporter now declares <|eot|> in the bundle's tokenizer config.
<|eom|> (200007) is deliberately not a stop token: it ends a message, not a
turn, and stopping there would cut the answer off inside the reasoning channel.
Generalization: a chat model whose turn-end token is not one of the five known
spellings will run to the budget on every request, and it looks like a verbose
model rather than a broken bundle. Check generation_config.json's eos_token_id
against what the bundle actually declares.
Quality: it is the lightest artifact and does not pay for it
The speed table above compares three artifacts at three different weights, and Core AI's is the smallest. That invites the obvious objection β some of the speed could just be fewer bytes. Measured, it isn't:
| weights | GSM8K, 100 questions | |
|---|---|---|
Core AI int4hu |
16.35 GB | 98 |
ExecuTorch k-quant-17G |
17.9 GB | 97 |
| MLX 4-bit | 18 GB | 95 |
Greedy, same 100 questions, scored by Yardstick's scripts/parity_gsm8k.py β the same
question set, CoT suffix, extractor and scoring the Gemma-4 campaign uses. That file was
not edited; the two arms it lacks (ExecuTorch via Meta's own solo_runner, MLX via
mlx_vlm because mlx_lm does not know muse_glimmer) are added around it.
Three questions apart is not a resolvable difference at n=100. The honest reading is "no arm is meaningfully worse", not "Core AI wins". Meta's published 1.0% degradation for this quant is below what 100 questions can see at all.
Two-pass budget, and why a one-pass number would have been wrong
llm-runner's wall time is 5 s + 0.037 x max_tokens on this bundle β independent of
how many tokens are actually generated. It keeps stepping to the budget after the stop
token halts output, so a budget wide enough for the longest answer taxes every question.
(ExecuTorch does not pay for unused budget.)
So pass 1 runs at 700 and pass 2 re-runs only the questions that hit it, at 2048. This is not just a speed trick β it changes the result. At budget 700 Core AI scored 87/100; 32 of those answers were truncated, and re-running them un-truncated took it to 98. A truncated answer is not blank, it is a wrong number the extractor picks up from the middle of the reasoning β and sometimes a right one by accident. One-pass at 700 would have published 87 and called it accuracy.
Truncation rates differ per arm (Core AI 32, MLX 26, ExecuTorch 14), so a single fixed budget would have penalised the arms that reason longer β a quality table measuring verbosity.
Residue, stated: 3 questions (Core AI), 2 (MLX), 1 (ExecuTorch) still hit 2048 and are scored from truncated output. Core AI and MLX also generate noticeably longer answers than ExecuTorch for the same questions (median 532 / 453 / 326 tokens); that is unexplained and does not show up in the score.
Gates
| stage | result |
|---|---|
Authoring vs HF transformers 5.15, fp32, real weights, 8 layers (2 NoPE) |
cos 1.0000 at embedding, every layer, final norm and logits; top-1 match; 0 argmax flips |
| Full-size key binding (52 layers) | 471/471 parameters bound; 0 missing / 0 unexpected / 0 shape mismatch |
This bundle vs the fp16 oracle, coreai_gate.py |
PASS β token-for-token over 24 generated tokens |
Run it
llm-benchmark --model gpu-pipelined/muse_glimmer_30b_decode_int4hu_block32_sym \
-p 512 -g 1024 -n 3
llm-runner --model gpu-pipelined/muse_glimmer_30b_decode_int4hu_block32_sym \
--prompt "..." --temperature 0.0 \
--inference-engine-variant coreai-pipelined --warmup off
Needs ~17 GB of unified memory for the weights plus the KV cache (52 KB/token; 0.44 GB at the 8192 context this bundle is exported for). Mac only β no quantization brings 16.35 GB to an iPhone.
Layout
gpu-pipelined/muse_glimmer_30b_decode_int4hu_block32_sym/
muse_glimmer_30b_decode_int4hu_block32_sym.aimodel/
metadata.json
tokenizer/ (incl. chat_template.jinja, copied verbatim)
config.json the source config, so zoo_verify can compare against it
LICENSE carried from the source repo (Apache-2.0)
Recipe, export script and the port's findings:
coreai-model-zoo β
models/muse-glimmer-30b/, conversion/export_muse_glimmer_decode_pipelined.py,
knowledge/muse-glimmer-port.md.
- Downloads last month
- 304
Model tree for mlboydaisuke/Muse-Glimmer-30B-CoreAI
Base model
meta-models/Muse-Glimmer-30B