Phonon-2 · experimental Core AI bundles

Three experimental conversions of Fermion Research's Phonon-2 for Apple silicon and public Core AI APIs on macOS/iOS 27. LUT6 is recommended for speed. The other variants expose the tradeoff between download size, working memory, and inference time.

Phonon-2 is an encoder-heavy FastConformer/TDT model derived from NVIDIA Parakeet V3. The conversions preserve its trained five-value encoder weights and original INT6 decoder values. They use FP16 activations and the publisher's pause-based chunking policy: 25–35 seconds, targeting 30 seconds, preserving all audio. English is the qualified language.

Directory Encoder storage Unpacked files LZRAVEN archive Qualification
lut6/ Exact six-bit palettes shared across eight rows 583.4 MB 283.2 MB M3 Mac and iPhone 15 Pro Max
lut4/ Exact four-bit palettes shared across two rows 438.6 MB 235.0 MB M3 Mac and iPhone 15 Pro Max
size/ Original packed encoder records, expanded into FP16 tensor inputs at load 217.2 MB 178.6 MB M3 Mac only

Sizes use decimal MB. Each archive restores its corresponding directory byte-for-byte. Core AI specializes the extracted .aimodel assets on the device; the archives contain portable source assets, not device caches. Each directory includes its own metadata.json, runtime.json, vocabulary, frontend constants, compact decoder data, license, and attribution.

Measured performance

JFK's “We choose to go to the Moon” speech: a 20-second excerpt and the full 18-minute 15-second recording. Measurements were collected September 30, 2026. Larger RTFx means faster. All Core AI values are three-run medians after preparation and warmup.

M3 MacBook Air, 16 GB, macOS 27.0.1. Includes frontend, encoder and decoding; excludes loading, file reading and chunk planning. Full-file planning adds about 0.06 seconds.

Variant 20-second excerpt Full speech Full RTFx Sampled full-pass memory
LUT6 0.251 s 2.461 s 445.0× 1.05–1.20 GB
LUT4 0.330 s 4.470 s 245.0× 0.90 GB
Size-focused 0.365 s 5.335 s 205.3× 2.18 GB
Publisher fermion MLX defaults 0.231 s 12.668 s 86.5× 11.34 GB

All memory figures use sampled process footprint plus simultaneous attributed accelerator no-footprint memory; external or unattributed driver/compiler memory is excluded. The publisher's 11.34 GB was measured during a dedicated timed API run. Its unmodified CLI reached 12.09 GB across loading, warmup and inference. Much of that was reusable allocation cache: clearing it after completion reduced the CLI footprint to 2.30 GB with active tensors unchanged. No lower-cache performance claim is implied.

The unmodified publisher runtime was fermion-research 0.2.4 with tdt16,dense16. Its timing includes chunk planning. A separate invocation of its own benchmark CLI confirmed a 12.537-second full-speech median (87.4×). The short excerpt favors the publisher runtime; these speedups describe the long-recording workload.

iPhone 15 Pro Max, A17 Pro, iOS 27.0.1, Release build. Includes bounded WAV reading, chunk planning, frontend, encoder and decoding; preparation and report serialization are excluded. Foreground, nominal thermals, Low Power Mode off.

Variant 20-second excerpt Full speech Full RTFx Sampled full-pass memory
LUT6 0.321 s 3.075 s 356.2× 1.03 GB
LUT4 0.370 s 5.349 s 204.8× 0.90 GB
Size-focused — — — Not measured on iPhone

Preparation and placement

Variant Observed first preparation, M3 Cached preparation, M3 Observed first preparation, iPhone
LUT6 104.2 s 0.35 s 72.6 s
LUT4 179.1 s 0.18–0.66 s 243.1 s
Size-focused 49.7 s 3.56–3.59 s Not measured

These are observed first-use sessions, not repeated clean-install medians. The LUT4 Mac run reused decoder caches. LUT6 and size-focused runs reused unchanged subsampling and decoder specializations. The size-focused variant expands about 1.21 GB of FP16 encoder tensors at load, accounting for most of its cached preparation time and extra working memory. Its initial preparation was measured before a benchmark-reporting issue was corrected; subsequent clean inference runs established its numerical and performance results.

Complete warmed traces for each qualified platform/variant show 493 ANE predictions and zero target-process GPU intervals: 259 predictions during subsampling/encoding and 234 during batched decoding. CPU feature extraction and host TDT control remain outside these graphs. Placement is evidence of ANE execution, not a measurement of arithmetic utilization.

Accuracy and runtime contract

All three conversions produce identical token IDs and transcript text on the two reference clips. The phone-qualified variants match their Mac outputs. Full-speech WER is 2.93% using Whisper's English normalizer (65 errors in 2,220 reference words), equal to an independent FP32 reconstruction's normalized word sequence; the publisher MLX defaults score 3.02%. This is one recording, not a general accuracy ranking. Two full-speech chunks differ in tokenization from FP32 while preserving the normalized words.

Native word timings are returned. They are finite, ordered and within each chunk; human word-boundary accuracy has not been measured independently. The size-focused variant shifts eight full-speech word intervals relative to the palette variants, by at most 160 ms, without changing the transcript.

Use a runtime that implements the included metadata and scheduling policy. The six encoder graphs each evaluate four layers over the complete 448-position context. Valid-length masks, relative-position conventions, silence handling, TDT state and duration decoding, and compact weight expansion are part of the runtime contract. runtime.json specifies four concurrent encoder requests and ANE decoding of up to 128 independent audio chunks. The size-focused variant uses a single batched decoder asset for both single utterances and batches; its encoder weights are explicit tensor inputs.

Download

Download one variant. On macOS 27, the smaller archive can be extracted with Apple's aa tool:

hf download coder543/phonon-2-coreai-experimental archives/phonon-2-lut6.aar --local-dir phonon-download
mkdir -p phonon-2-lut6
aa extract -i phonon-download/archives/phonon-2-lut6.aar -d phonon-2-lut6

For individual runtime files:

hf download coder543/phonon-2-coreai-experimental --include 'lut6/*' --local-dir phonon-download

Choose lut4 or size in those paths for another design. Applications using compressed delivery need an AppleArchive/LZRAVEN extraction step before Core AI loading. Hugging Face supplies file checksums through its snapshot metadata.

Attribution

Weights are CC-BY-4.0, derived from NVIDIA Parakeet TDT 0.6B V3 and retrained/quantized by Fermion Research. The original license and NOTICE are included. Conversion changes include staged Core AI graphs, FP16 activations, exact palette or external-weight storage, compact decoder packing, and batched decoding. No retraining or pruning is performed. Compiler source-location metadata is removed from the published assets.

Source revision: 9c7fef3584499a88fe8d394427f45851bbb8b446. Tokenizer/configuration revision: 541d1f99c6b0c3cd0b11a95167540bb8edefd82b. These identify conversion provenance; consumers should resolve a consistent current repository snapshot and honor its runtime metadata.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for coder543/phonon-2-coreai-experimental

Quantized
(3)
this model

Collection including coder543/phonon-2-coreai-experimental