short-reply-0.6b-coreml
A sub-1B model that reads one social post and drafts one short reply (≤ 12 words), packaged for Core ML. Qwen3-0.6B, LoRA-tuned by distilling a 1.7B reply model on 4,615 generated posts, then merged. One 724 MB package with two functions over shared weights:
| Function | Shape | Runs on | p50 (M5 Pro, macOS 27, from Swift) |
|---|---|---|---|
prefill |
160 tokens, fixed | Neural Engine (1,920 / 1,921 ops) | 28 ms |
decode |
1 token, 512-slot KV cache in MLState |
GPU | 24 ms / token |
A reply takes about 190 ms. Weights are 8-bit palettized (k-means, per-grouped-channel, group 32); the shared
embedding / lm_head weight stays fp16 (a palettized embedding table costs +13 ms per token on the gather).
Linear int8 weights run on the ANE but return garbage there; palettized weights work.
Quality
On 100 held-out real posts, 89 replies are identical to the PyTorch checkpoint; a blind LLM-judge A/B against the project's 1.7B model scores 88 vs 93 (all of relevant / tone / grounded / ≤ 12 words), the same as the PyTorch 0.6B. About 1 in 8 drafts is generic or slightly off. Replies are drafts for a person to review, not for posting automatically. Training used only generated and hand-written posts; no real social posts were used for training.
Files
short_reply_0_6b.mlpackage— multifunction Core ML package (prefill,decode), token-id inputs (embedding lookup and RoPE tables are inside the graph).tokenizer.json— Qwen3 byte-level BPE.config.json— host settings: prefill length 160, cache 512, pad / stop ids, max new tokens 32, and the system prompt the model was trained with.
Demo
Menu-bar app in FluidInference/FluidUse (ShortReplyDemo):
open a post's reply box in any app, press 9, the draft is pasted into the box; 0 regenerates it. About
190 ms per reply with the Neural Engine on the prompt and the GPU on the tokens.
hf download FluidInference/short-reply-0.6b-coreml --local-dir ~/Models/short-reply-0.6b-coreml
Sources/ShortReplyDemo/demo.sh --x ~/Models/short-reply-0.6b-coreml
Use
Swift host: ShortReplyManager in the same repository.
let replies = try await ShortReplyManager.load(from: modelDirectory)
let draft = try await replies.draft(for: "Finally passed my driving test on the third try.")
print(draft.reply, draft.timing.totalSeconds)
Prompt (Qwen3 chat template, thinking disabled):
<|im_start|>system
{config.systemPrompt}<|im_end|>
<|im_start|>user
Post: {post}
Reply:<|im_end|>
<|im_start|>assistant
<think>
</think>
Left-pad the token ids to 160 with padID, mask padded columns, run prefill, copy its keys / values
([28, 8, 160, 128] fp16) into the decoder state (k_cache_i / v_cache_i, [1, 8, 512, 128]), then call
decode with one token at a time (input_ids, attention_mask [1, 1, 1, 512], cache_position) until
<|im_end|> or 32 tokens. Greedy decoding reproduces the benchmarked replies.
Provenance
Base: Qwen/Qwen3-0.6B (Apache-2.0). Conversion: coremltools 9.0,
minimum_deployment_target macOS 15 / iOS 18. Training data: 4,615 generated posts labeled by a 1.7B reply model
plus 478 hand-written pairs; no real social posts. Training and conversion records are kept in the author's lab
repository.
- Downloads last month
- 380