clef-flash (Core ML)
Cloudflare's clef-flash, a 9B decision model post-trained from Qwen3.5-9B, converted to Core ML for Apple silicon. Text only, 8-bit weights, runs on the Mac GPU.
Same contract as Clef and the Jev / SystemOne API: a state (text or JSON) and typed questions (noul yes/no,
choice, ordered score) go in; one probability for every allowed option of every question comes out of a
single forward pass. Nothing is generated.
Runtime: ClefFlashManager in FluidUse (Swift, macOS 15+), with a
SwiftUI support-ticket triage demo (ClefFlashDemo).
Packages
| File | Role | Size | Precision |
|---|---|---|---|
part00 โฆ part07.mlpackage |
Qwen3.5 decoder, 4 layers each (the last adds the final norm); functions L256 L512 L1024 L2048 per sequence bucket, weights shared |
6.5 GB total | 8-bit weights (per channel), fp16 compute |
Head.mlpackage |
Clef's joint schema head, fixed shape: โค 16 questions, โค 96 options; same four functions | 465 MB | fp32 |
embeddings.f16 |
input token embeddings (host-side gather) | 2.0 GB | fp16 |
output_embeddings.f16 |
output embeddings, the head's lexical option rows (clef-flash does not tie them) | 2.0 GB | fp16 |
tokenizer.json, config.json |
tokenizer; shapes, buckets and token ids for the host |
Each decoder part: hidden [1, L, 4096], cos / sin [L, 64] (fp16) โ out [1, L, 4096]. Text-only records use
positions 0..L-1 on all three M-RoPE axes. The host renders the record exactly like Clef's encode_record,
gathers embeddings, chains the 8 parts and builds the head's span-mean matrices; the Swift runtime is the reference
host. Images and video are not converted.
Quality
Reference: the same decoder in fp32 (matches Hugging Face's Qwen3_5TextModel to 5e-5 relative) streamed four
layers at a time, plus Cloudflare's own JointSchemaHead. 605 records / 611 questions (README samples, ARC-Easy and
ARC-Challenge test, first 300 each):
| This model (8-bit Core ML) | fp32 reference | Cloudflare published | |
|---|---|---|---|
| answers that differ from fp32 | 1 / 611 (a near-tie: fp32 top-2 within 0.005) | โ | โ |
| median probability difference | 0.0004 | โ | โ |
| ARC-Easy | 100.0% | 100.0% | 99.5% |
| ARC-Challenge | 99.0% | 98.67% | 98.3% |
ARC rendering is Clef's record format with a plain state, not the Decision Index kit's templates, so compare within about a point.
Speed (M5 Pro, 24 GB, GPU)
| Median | |
|---|---|
| one decoder part (4 layers), 512-token bucket | 62 ms |
| one record, 3 questions, ~370 tokens (L512), Swift runtime | 0.49 s (p95 0.57 s) |
| load (compiled), 8 parts + head | ~40 s |
The decoder needs ~7 GB of memory to itself; with other large processes running it pages and slows down sharply. The Qwen3.5 decoder (Gated DeltaNet layers) does not run on the Neural Engine: Core ML places 0 of its ops on the ANE, so this is a GPU model.
Usage (Swift)
import FluidUse
let manager = try await ClefFlashManager.load(from: bundleDirectory, bucket: 512)
let result = try await manager.answer(
state: "Our checkout is down and customers can't pay.",
questions: [
("team", .choice(instructions: "Which team should handle this ticket?",
criteria: ["billing": "Charges, refunds", "engineering": "Bugs, outages"])),
("urgency", .score(instructions: "How urgent is this ticket?", criteria: ["Low", "Normal", "High", "Critical"])),
])
print(result.answers.map { ($0.questionID, $0.choice) }, result.totalMilliseconds)
License and credits
Apache-2.0, same as Cloudflare/clef-flash (LICENSE is
Cloudflare's). All model weights are Cloudflare's; this repository only changes their format and precision.
Conversion by Fluid Inference.
- Downloads last month
- 88