short-reply-0.6b-coreml

A sub-1B model that reads one social post and drafts one short reply (≤ 12 words), packaged for Core ML. Qwen3-0.6B, LoRA-tuned by distilling a 1.7B reply model on 4,615 generated posts, then merged. One 724 MB package with two functions over shared weights:

Function Shape Runs on p50 (M5 Pro, macOS 27, from Swift)
prefill 160 tokens, fixed Neural Engine (1,920 / 1,921 ops) 28 ms
decode 1 token, 512-slot KV cache in MLState GPU 24 ms / token

A reply takes about 190 ms. Weights are 8-bit palettized (k-means, per-grouped-channel, group 32); the shared embedding / lm_head weight stays fp16 (a palettized embedding table costs +13 ms per token on the gather). Linear int8 weights run on the ANE but return garbage there; palettized weights work.

Quality

On 100 held-out real posts, 89 replies are identical to the PyTorch checkpoint; a blind LLM-judge A/B against the project's 1.7B model scores 88 vs 93 (all of relevant / tone / grounded / ≤ 12 words), the same as the PyTorch 0.6B. About 1 in 8 drafts is generic or slightly off. Replies are drafts for a person to review, not for posting automatically. Training used only generated and hand-written posts; no real social posts were used for training.

Files

  • short_reply_0_6b.mlpackage — multifunction Core ML package (prefill, decode), token-id inputs (embedding lookup and RoPE tables are inside the graph).
  • tokenizer.json — Qwen3 byte-level BPE.
  • config.json — host settings: prefill length 160, cache 512, pad / stop ids, max new tokens 32, and the system prompt the model was trained with.

Demo

Menu-bar app in FluidInference/FluidUse (ShortReplyDemo): open a post's reply box in any app, press 9, the draft is pasted into the box; 0 regenerates it. About 190 ms per reply with the Neural Engine on the prompt and the GPU on the tokens.

hf download FluidInference/short-reply-0.6b-coreml --local-dir ~/Models/short-reply-0.6b-coreml
Sources/ShortReplyDemo/demo.sh --x ~/Models/short-reply-0.6b-coreml

Use

Swift host: ShortReplyManager in the same repository.

let replies = try await ShortReplyManager.load(from: modelDirectory)
let draft = try await replies.draft(for: "Finally passed my driving test on the third try.")
print(draft.reply, draft.timing.totalSeconds)

Prompt (Qwen3 chat template, thinking disabled):

<|im_start|>system
{config.systemPrompt}<|im_end|>
<|im_start|>user
Post: {post}
Reply:<|im_end|>
<|im_start|>assistant
<think>

</think>

Left-pad the token ids to 160 with padID, mask padded columns, run prefill, copy its keys / values ([28, 8, 160, 128] fp16) into the decoder state (k_cache_i / v_cache_i, [1, 8, 512, 128]), then call decode with one token at a time (input_ids, attention_mask [1, 1, 1, 512], cache_position) until <|im_end|> or 32 tokens. Greedy decoding reproduces the benchmarked replies.

Provenance

Base: Qwen/Qwen3-0.6B (Apache-2.0). Conversion: coremltools 9.0, minimum_deployment_target macOS 15 / iOS 18. Training data: 4,615 generated posts labeled by a 1.7B reply model plus 478 hand-written pairs; no real social posts. Training and conversion records are kept in the author's lab repository.

Downloads last month
380
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FluidInference/short-reply-0.6b-coreml

Finetuned
Qwen/Qwen3-0.6B
Finetuned
(1334)
this model