Core AI models for Privacy AI

Apple Core AI (.aimodel / .aimodelc) bundles for iOS 27 and macOS 27, exported from the original PyTorch weights with Apple's coreai-models toolchain and compiled ahead of time for the Apple Neural Engine.

These are used by Privacy AI. They are plain Core AI bundles, so anything that can load a Core AI model can use them.

Why Neural Engine builds

On iOS the GPU is revoked when an app is backgrounded; the Neural Engine is not. Measured on an iPhone 16 Pro Max (iPhone17,2, iOS 27.0), qwen3-0.6b-coreai:

prefill decode
foreground 84 t/s 76.5 t/s
backgrounded 1507 t/s 60.3 t/s

Backgrounded decode runs at 79% of foreground. That is what makes these usable from a keyboard extension and from background assistant work.

Layout

Each leaf directory is a complete, self-contained bundle (metadata.json, tokenizer/, one model asset). Download exactly one.

ios/<model-id>/<architecture>/   ahead-of-time compiled for one silicon generation
ios/<model-id>/source/           architecture-neutral .aimodel, runs everywhere
macos/<model-id>/                macOS build (dynamic shapes, GPU)

Pick the subdirectory matching AIModel.deviceArchitectureName at runtime, and fall back to source/ when there is no prebuilt bundle for that architecture. source/ works on every chip including unreleased ones, at the cost of a several-minute first-load specialization; a matching prebuilt architecture loads in well under a second once its cache is warm.

architecture chips
h17p A18, A18 Pro
h18p A19, A19 Pro

Models

id kind leaf size context structure architectures upstream
qwen3-0.6b-coreai chat 554 MB 4096 chunked-static (ANE) h17p, h18p, source Qwen/Qwen3-0.6B
qwen3-0.6b-coreai (macOS) chat 331 MB 8192 dynamic (GPU) — Qwen/Qwen3-0.6B

Reproduce the iOS build with:

cd coreai-models
uv run coreai.llm.export qwen3-0.6b --platform iOS --max-context-length 4096 --output-dir out
xcrun coreai-build compile out/qwen3_0_6b_mixed_4bit_8bit_static/qwen3_0_6b_mixed_4bit_8bit_static.aimodel \
    --output h17p/qwen3_0_6b_mixed_4bit_8bit_static.aimodelc \
    --platform iOS --min-deployment-version 27.0 \
    --preferred-compute neural-engine --architecture h17p

Note on context length

iOS Core AI exports use static shapes, so the context is fixed at export time and its KV cache is resident whenever the model is loaded. It is also not a free choice: the ANE compiler refuses a graph whose largest KV tensor exceeds roughly 2.1x10^8 elements, and coreai-build compile returns 0 and writes a complete-looking .aimodelc even when ANE compilation failed entirely, leaving a GPU-only bundle behind.

The only reliable check is that the compiled asset contains one ANE bitcode per region:

find <bundle>.aimodelc -name '*.mlir.bc' | wc -l   # expect tiers x 3 x 2

Every bundle published here passes that check (30/30 regions for the 4096-context builds).

Licence and attribution

The models are redistributed under their original licences, unchanged. qwen3-0.6b-coreai derives from Qwen/Qwen3-0.6B by Alibaba Cloud, Apache License 2.0 (see LICENSE). Conversion to Core AI format was done with Apple's coreai-models export recipes; no weights were retrained or modified beyond the quantization the export applies.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support