Clef-Flash-ANE-LUT6

Model architecture

A mixed-precision conversion of Cloudflare/clef-flash for the Apple Neural Engine (ANE) on Apple Silicon Macs. It answers questions about text, images and short videos with yes/no, choice and score outputs; it does not generate text.

Component Structure
Text backbone 32 layers; hidden size 4,096; MLP intermediate size 12,288
Attention layout 24 Gated DeltaNet linear-attention layers (16 key and 32 value heads, head size 128) and 8 full-attention layers (16 query and 4 key-value heads, head size 256), with full attention every fourth layer
Decision head 4 layers; width 1,024; 16 heads; 2 routing layers
Vocabulary 248,320 tokens
Vision encoder 27 layers; hidden size 1,152; 16 heads; 16ร—16 patches with two-frame temporal patches; 2ร—2 merge into the 4,096-wide text embedding space. Shared by images and video
Context Up to 16,384 tokens, processed 128 tokens per ANE call; full-attention programs for 1,024, 2,048, 4,096, 8,192 and 16,384 tokens

The payload is approximately 9.74 GiB: ANE programs 4.86 GiB, token embeddings and LM head 3.79 GiB, decision head 0.23 GiB, vision encoder 0.85 GiB. These are on-disk sizes, not RAM requirements.

Component Runs on
Text backbone (32 layers) ANE
Vision transformer (27 layers) ANE, FP16. One fixed program per layer over 1,024 patches; images and video frame pairs are packed with a block-diagonal attention mask
Patch embedding, vision positions and 2ร—2 merger CPU
Token embedding lookup and decision head CPU

No GPU is used.

Format compatibility: this is a custom artifact, not a standard Transformers, MLX or Core ML .mlpackage checkpoint. The text layers are stored as MIL program descriptions with packed weights, compiled on the destination Mac when loaded. Other tensors are safetensors. Loading uses Apple's private ANE interface, which may change with macOS updates; only an M5 Pro running macOS 27 has been tested.

Weight precision

Component Stored precision
Large text linear layers: MLP gate/up/down, linear-attention QKV/Z/output, full-attention Q/K/V/O LUT-6: 6-bit indices with one 64-entry FP16 table per 16 output channels
Other text-program weights: normalization, short convolution, small gate projections and decay parameters FP16, converted from BF16
Token embeddings, LM head and final normalization Original BF16
Decision head Original BF16
Vision encoder Original BF16; converted to FP16 when loaded onto the ANE. Not LUT-quantized

LUT-6 text weights

  • Every linear layer with at least 1,048,576 weights is quantized. Each group of 16 output channels has its own 64-value table, found by 40 rounds of one-dimensional k-means. Half of the starting values are quantiles; the other half are spread evenly between the group's minimum and maximum, so that outlying weights also get table values.
  • Four 6-bit indices are packed into three bytes in Core ML's little-endian layout. Including the tables, a layer with 4,096 inputs stores about 6.02 bits per weight.
  • Activations run in FP16 on the ANE. The MLP down-projection input is scaled up by as much as 16ร— and attention probabilities by 16ร— to keep precision in FP16; both are scaled back afterwards.
  • Quantization is lossy; the original BF16 weights cannot be reconstructed exactly.

Vision weights

  • LUT-6 was tested on the vision layers and rejected: per-layer error was 5โ€“14ร— higher than FP16, with only about 10% speedup.
  • In FP16 on the ANE, per-layer relative RMSE against float32 was 0.08โ€“0.22%, compared with 0.30โ€“0.55% for BF16 on the Metal GPU.
  • The largest residual activation observed was about 12,400, within the FP16 limit of 65,504. Inputs that overflow are rejected rather than answered.
  • The vision programs are built from these weights on the destination Mac when the first image or video arrives, which took about 10 seconds on an M5 Pro.

The original tokenizer, chat template, configuration and image/video processor settings are retained. File sizes and SHA-256 checksums are recorded in the bundle manifest.

Validation

On a fixed regression set of 17 synthetic text, image and video requests (27 questions), the original BF16 multimodal model was the reference, with identical preprocessing:

  • The selected answer matched on 27 of 27 questions; the largest probability difference was 0.065.
  • Both models met 26 of the 27 synthetic expectations. The shared miss is a mixed request in which a separate green image is read as the start of a red-to-blue video.
  • During those requests, repeated twice, the serving process recorded 0 ns of GPU time on an M5 Pro running macOS 27.

This is a regression check, not a general accuracy benchmark. Images were resized with Pillow to at most 262,144 pixels (256 visual tokens). Video was sampled about twice per second by presentation timestamp, up to 64 frames of at most 65,536 pixels each. Audio is not used.

The earlier text-only conversion is available at revision text-v1.

Source and license

Source revision: 17f0b0ad64efb65d273590632833508766b2aae6.

Original model Clef-Flash by Cloudflare, post-trained from Qwen3.5-9B by Qwen. Conversion by Yanun. Released under Apache-2.0, the license of the source model. Redistributions must retain LICENSE and NOTICE.

This is not an official Cloudflare, Qwen or Apple release.

Downloads last month
22
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for Yanun/clef-flash-ane-lut6

Finetuned
Qwen/Qwen3.5-9B
Quantized
(45)
this model