Clef-Flash-ANE-LUT6
Model architecture
A mixed-precision conversion of Cloudflare/clef-flash for the Apple Neural Engine (ANE) on Apple Silicon Macs. It answers questions about text, images and short videos with yes/no, choice and score outputs; it does not generate text.
| Component | Structure |
|---|---|
| Text backbone | 32 layers; hidden size 4,096; MLP intermediate size 12,288 |
| Attention layout | 24 Gated DeltaNet linear-attention layers (16 key and 32 value heads, head size 128) and 8 full-attention layers (16 query and 4 key-value heads, head size 256), with full attention every fourth layer |
| Decision head | 4 layers; width 1,024; 16 heads; 2 routing layers |
| Vocabulary | 248,320 tokens |
| Vision encoder | 27 layers; hidden size 1,152; 16 heads; 16ร16 patches with two-frame temporal patches; 2ร2 merge into the 4,096-wide text embedding space. Shared by images and video |
| Context | Up to 16,384 tokens, processed 128 tokens per ANE call; full-attention programs for 1,024, 2,048, 4,096, 8,192 and 16,384 tokens |
The payload is approximately 9.74 GiB: ANE programs 4.86 GiB, token embeddings and LM head 3.79 GiB, decision head 0.23 GiB, vision encoder 0.85 GiB. These are on-disk sizes, not RAM requirements.
| Component | Runs on |
|---|---|
| Text backbone (32 layers) | ANE |
| Vision transformer (27 layers) | ANE, FP16. One fixed program per layer over 1,024 patches; images and video frame pairs are packed with a block-diagonal attention mask |
| Patch embedding, vision positions and 2ร2 merger | CPU |
| Token embedding lookup and decision head | CPU |
No GPU is used.
Format compatibility: this is a custom artifact, not a standard Transformers, MLX or Core ML .mlpackage checkpoint. The text layers are stored as MIL program descriptions with packed weights, compiled on the destination Mac when loaded. Other tensors are safetensors. Loading uses Apple's private ANE interface, which may change with macOS updates; only an M5 Pro running macOS 27 has been tested.
Weight precision
| Component | Stored precision |
|---|---|
| Large text linear layers: MLP gate/up/down, linear-attention QKV/Z/output, full-attention Q/K/V/O | LUT-6: 6-bit indices with one 64-entry FP16 table per 16 output channels |
| Other text-program weights: normalization, short convolution, small gate projections and decay parameters | FP16, converted from BF16 |
| Token embeddings, LM head and final normalization | Original BF16 |
| Decision head | Original BF16 |
| Vision encoder | Original BF16; converted to FP16 when loaded onto the ANE. Not LUT-quantized |
LUT-6 text weights
- Every linear layer with at least 1,048,576 weights is quantized. Each group of 16 output channels has its own 64-value table, found by 40 rounds of one-dimensional k-means. Half of the starting values are quantiles; the other half are spread evenly between the group's minimum and maximum, so that outlying weights also get table values.
- Four 6-bit indices are packed into three bytes in Core ML's little-endian layout. Including the tables, a layer with 4,096 inputs stores about 6.02 bits per weight.
- Activations run in FP16 on the ANE. The MLP down-projection input is scaled up by as much as 16ร and attention probabilities by 16ร to keep precision in FP16; both are scaled back afterwards.
- Quantization is lossy; the original BF16 weights cannot be reconstructed exactly.
Vision weights
- LUT-6 was tested on the vision layers and rejected: per-layer error was 5โ14ร higher than FP16, with only about 10% speedup.
- In FP16 on the ANE, per-layer relative RMSE against float32 was 0.08โ0.22%, compared with 0.30โ0.55% for BF16 on the Metal GPU.
- The largest residual activation observed was about 12,400, within the FP16 limit of 65,504. Inputs that overflow are rejected rather than answered.
- The vision programs are built from these weights on the destination Mac when the first image or video arrives, which took about 10 seconds on an M5 Pro.
The original tokenizer, chat template, configuration and image/video processor settings are retained. File sizes and SHA-256 checksums are recorded in the bundle manifest.
Validation
On a fixed regression set of 17 synthetic text, image and video requests (27 questions), the original BF16 multimodal model was the reference, with identical preprocessing:
- The selected answer matched on 27 of 27 questions; the largest probability difference was 0.065.
- Both models met 26 of the 27 synthetic expectations. The shared miss is a mixed request in which a separate green image is read as the start of a red-to-blue video.
- During those requests, repeated twice, the serving process recorded 0 ns of GPU time on an M5 Pro running macOS 27.
This is a regression check, not a general accuracy benchmark. Images were resized with Pillow to at most 262,144 pixels (256 visual tokens). Video was sampled about twice per second by presentation timestamp, up to 64 frames of at most 65,536 pixels each. Audio is not used.
The earlier text-only conversion is available at revision text-v1.
Source and license
Source revision: 17f0b0ad64efb65d273590632833508766b2aae6.
Original model Clef-Flash by Cloudflare, post-trained from Qwen3.5-9B by Qwen. Conversion by Yanun. Released under Apache-2.0, the license of the source model. Redistributions must retain LICENSE and NOTICE.
This is not an official Cloudflare, Qwen or Apple release.
- Downloads last month
- 22