Phonon-2 Core ML

Phonon-2 for the Apple Neural Engine. The Core ML package reads with the accuracy of the Phonon-2 reference engine: on the Open ASR Leaderboard's seven English test sets, run with the leaderboard's own code on the full sets, it averages 5.21 % word error, with LibriSpeech test-clean at 1.73 % and test-other at 3.91 %.

On an M5 MacBook Air an hour of speech becomes text in 6 seconds (606 times realtime), a few seconds of audio returns its text in about 11 milliseconds once the model is loaded, every word carries its start and end time, and the GPU stays free.

Run it

Phonon-2 is the model inside Detta, the dictation app for the Mac. This repository runs it on the Neural Engine from Swift or Python.

Download the model folder with the Hugging Face command line, in a virtual environment with Python 3.10 to 3.13 (macOS's built-in python3 is 3.9, and coremltools has no wheels for 3.14 yet):

python3.12 -m venv .venv && source .venv/bin/activate
pip install huggingface_hub
hf download FermionResearch/Phonon-2-CoreML --local-dir Phonon-2-CoreML

Swift, on macOS 15 or iOS 18 and later, with the package at github.com/fermionresearch/phonon-coreml (also in phonon-coreml-src.tgz). Build the command-line tool from the GitHub release and point it at the model folder:

git clone --branch 1.1.1 https://github.com/fermionresearch/phonon-coreml
(cd phonon-coreml && swift build -c release)
phonon-coreml/.build/release/phonon-coreml-cli Phonon-2-CoreML recording.m4a --words

It prints the transcript, then every word on its own line with its start and end time in seconds; --json words.json writes the same as JSON. From your own code:

import PhononCoreML

let transcriber = try Transcriber(bundle: URL(fileURLWithPath: "Phonon-2-CoreML"))   // the folder downloaded above
let result = try transcriber.transcribe(url: URL(fileURLWithPath: "recording.wav"))
print(result.text)
for word in result.words { print(word.start, word.end, word.text) }

Python, with coremltools, in the same virtual environment:

pip install coremltools numpy soundfile scipy
python Phonon-2-CoreML/phonon_coreml.py Phonon-2-CoreML recording.wav

The Swift runner reads wav, m4a, mp3, aiff, caf and flac at any sample rate and length; the Python runner reads wav and flac at any sample rate. The first run prepares the package for the Neural Engine once, which can take a minute or two; the compiled copy is kept, so later loads take under a second.

Phonon-2 also runs through the fermion command line on Apple silicon, Linux, Windows and NVIDIA GPUs. See the Phonon-2 card.

Long audio

An utterance of up to 35 seconds is read in one pass by the smallest encoder window that holds it. Longer recordings are cut at pauses into windows of up to 15 seconds and their words are joined, so no word is split at a boundary. Pass recordings whole; both runners do the cutting.

Accuracy

Every utterance of the seven English test sets, scored with the Open ASR Leaderboard's own code, with the Core ML package beside the Phonon-2 MLX engine.

Set Core ML, Neural Engine Phonon-2 MLX engine
LibriSpeech clean 1.73 1.72
LibriSpeech other 3.91 3.92
AMI 9.33 9.37
Earnings-22 6.99 6.96
GigaSpeech 8.33 8.35
SPGISpeech 3.70 3.70
VoxPopuli 2.46 2.46
Mean of seven 5.21 5.21

Word error rate, %, lower is better. Phonon-2 transcribes English. The Phonon-2 card lists its accuracy in other languages.

Speed

On an M5 MacBook Air, with the model loaded once, the encoder on the Neural Engine and the decoder on the CPU.

Audio Time Times real time
One hour of LibriSpeech read as a single file 6.0 s 606×
158 hours of test-set audio, utterance by utterance 31 min 305×
20 recordings of 2 seconds to 2 minutes (797 s) 2.0 s 400×
A recording of 3 to 5 seconds, model loaded (median of 20) 11 ms

Files

File Size Contents
Phonon-2.mlpackage 332 MB the encoder for 5, 10, 15 and 35 second windows over one shared weight set (five-value weights stored exactly, 16-bit float activations); audio is masked to its true length inside the window
decoder.bin 13 MB the prediction network, joint and vocabulary
manifest.json source checksum, window lengths and input contract
phonon_coreml.py the Python runner, with frontend.py and segmenter_np.py
phonon-coreml-src.tgz the Swift package and command-line tool
ci/ three LibriSpeech clips and their transcripts, for a quick check
SHA256SUMS a checksum for every file above

Licence

The weights are released under CC-BY-4.0. They derive from NVIDIA's parakeet-tdt-0.6b-v3, whose tokenizer and output conventions (punctuation, casing, numerals) they keep, and NOTICE lists the changes. The Swift package, the Python runner and this repository's code are released under Apache 2.0.

Downloads last month
37
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FermionResearch/Phonon-2-CoreML

Quantized
(6)
this model