WeMM-Embedding-2B CoreML (Apple Neural Engine)

A Core ML conversion of tencent/WeMM-Embedding-2B that runs on the Apple Neural Engine (ANE). It embeds text, images and short videos into the same 2,048-dimensional, L2-normalized vector space as the original model.

This repository holds model files only. To serve them, use embed-ane: a macOS menu bar app and CLI with an OpenAI-compatible POST /v1/embeddings endpoint. It downloads this repository and verifies every file digest before loading.

Files

Path What it is Size
chunks/chunk{0..5}.mlmodelc The 24 decoder layers as six 4-layer Core ML programs (fp16, sequence 512); rotary positions are inputs 6 × ~462 MB
vision_tower.mlmodelc Vision encoder and merger as one Core ML program (fp16); one call per image or video frame-pair ~628 MB
embed_table.fp16.npy Token embedding table, looked up on the CPU ~1.0 GB
vision_abs_pos.fp32.npy Learned vision position table, interpolated on the CPU 9.4 MB
tokenizer.json, tokenizer_config.json Unchanged from the original model 20 MB
spec.yaml Digest-pinned file list read by the embed-ane downloader 3 KB

Every .mlmodelc directory contains manifest.txt, a list of path size sha256 for its members.

Usage

With embed-ane installed, download and verify the model:

embed-ane fetch maolon/WeMM-Embedding-2B-CoreML-ANE
embed-ane serve

Then call the endpoint like any OpenAI embeddings API:

curl http://127.0.0.1:8080/v1/embeddings -H 'Content-Type: application/json' \
  -d '{"model": "wemm-embedding-2b-ane", "input": ["How is mapo tofu prepared?"]}'

Images and videos use content parts with local file URLs:

{"input": [{"type": "video_url", "video_url": {"url": "file:///path/to/clip.mp4"}},
           {"type": "text", "text": "optional text embedded together with the video"}]}

The menu bar app does the same from Models → Add Model.

How it differs from the original

The weights are the original WeMM-Embedding-2B weights converted to fp16. Nothing was retrained or distilled. Changes:

  • Precision. Core ML programs use fp16 storage and compute (compute_precision=FLOAT16), with a macOS 15 / iOS 18 deployment target.
  • Decoder split. The 24 Qwen3.5 decoder layers are split into six 4-layer chunks with a fixed 512-token sequence. The gated delta-rule (linear attention) layers are rewritten with ANE-friendly operations (cumulative-sum matmul, convolution shift and a block solver), which are equivalent in fp32.
  • Positions as inputs. Multimodal rotary positions (mRoPE) are computed on the CPU and passed to every chunk as cos/sin. Text, images and video share one decoder.
  • Tables off the graph. The token embedding table and the vision position table ship as .npy files and are applied on the CPU.
  • Vision tower. The vision encoder (24 blocks plus merger) is a single graph with a 1,536-patch (384-token) capacity. Its patch embedding takes a frame pair, as in the original, so stills (one frame repeated) and video use the same program.

Accuracy

Cosine similarity between this conversion, served by embed-ane on an M1 Pro ANE, and the original model in fp32 with the upstream prompt format:

Input Cosine to fp32 original
Text (6 sentences, zh/en) 0.977 mean (0.966–0.990)
Image + text (672×448) 0.949
Video (57 s clip, 8 frames) 0.920

Prompt token counts equal the original processor's in every case. Most of the gap is fp16 rounding in the decoder, which grows with sequence length (the image and video prompts are 306 and 488 tokens). Method and details: embed-ane docs/accuracy.md.

Speed on an M1 Pro: about 1.5 s per text, 2.5 s per image, 3.9 s per 8-frame video, with a 2.0 GB memory footprint.

Limits

  • Context. The context is 512 tokens. Text inputs may use up to 507 tokens; an image uses at most 384.
  • Video. Videos are sampled at 2 fps with at least 4 and at most 8 frames, uniformly over the clip. Resolution is chosen so all frame-pairs, timestamps and text fit in 512 tokens. The original model allows up to 768 frames at higher resolution, so long or detailed videos are represented more coarsely here.
  • One media item per input. Each input holds one image or one video plus text.
  • First load. The first load on a Mac compiles the programs for the ANE, which can take several minutes. macOS caches the result. If the startup disk is nearly full, macOS may purge that cache, and the next load recompiles.
  • Hardware. Apple silicon with macOS 15 or later.

License and attribution

The original model is © Tencent and licensed under Apache-2.0, with third-party components (Qwen3.5-2B) under their own licenses. See LICENSE, copied unchanged from the original repository. This conversion is distributed under the same terms. The modifications are listed above.

If you use this model, please cite the original work:

@article{wemm-embedding,
  title={WeMM-Embedding: WeChat Multi-Modal Embedding Technical Report},
  author={Junjie Zhou and Ke Mei and Lei Li and Tianyi Wang and Fengyun Rao and Jing Lyu},
  year={2026},
  eprint={2608.24053},
  archivePrefix={arXiv},
  primaryClass={cs.CV},
  url={https://arxiv.org/abs/2608.24053}
}
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for maolon/WeMM-Embedding-2B-CoreML-ANE

Finetuned
Qwen/Qwen3.5-2B
Quantized
(8)
this model

Paper for maolon/WeMM-Embedding-2B-CoreML-ANE