Phonon-2 Core ML
Phonon-2 for the Apple Neural Engine. The Core ML package reads with the accuracy of the Phonon-2 reference engine: on the Open ASR Leaderboard's seven English test sets, run with the leaderboard's own code on the full sets, it averages 5.21 % word error, with LibriSpeech test-clean at 1.73 % and test-other at 3.91 %.
On an M5 MacBook Air an hour of speech becomes text in 6 seconds (606 times realtime), a few seconds of audio returns its text in about 11 milliseconds once the model is loaded, every word carries its start and end time, and the GPU stays free.
Run it
Phonon-2 is the model inside Detta, the dictation app for the Mac. This repository runs it on the Neural Engine from Swift or Python.
Download the model folder with the Hugging Face command line, in a virtual environment with Python 3.10 to 3.13 (macOS's built-in python3 is 3.9, and coremltools has no wheels for 3.14 yet):
python3.12 -m venv .venv && source .venv/bin/activate
pip install huggingface_hub
hf download FermionResearch/Phonon-2-CoreML --local-dir Phonon-2-CoreML
Swift, on macOS 15 or iOS 18 and later, with the package at github.com/fermionresearch/phonon-coreml
(also in phonon-coreml-src.tgz). Build the command-line tool from the GitHub release and point it at the model folder:
git clone --branch 1.1.1 https://github.com/fermionresearch/phonon-coreml
(cd phonon-coreml && swift build -c release)
phonon-coreml/.build/release/phonon-coreml-cli Phonon-2-CoreML recording.m4a --words
It prints the transcript, then every word on its own line with its start and end time in seconds; --json words.json writes
the same as JSON. From your own code:
import PhononCoreML
let transcriber = try Transcriber(bundle: URL(fileURLWithPath: "Phonon-2-CoreML")) // the folder downloaded above
let result = try transcriber.transcribe(url: URL(fileURLWithPath: "recording.wav"))
print(result.text)
for word in result.words { print(word.start, word.end, word.text) }
Python, with coremltools, in the same virtual environment:
pip install coremltools numpy soundfile scipy
python Phonon-2-CoreML/phonon_coreml.py Phonon-2-CoreML recording.wav
The Swift runner reads wav, m4a, mp3, aiff, caf and flac at any sample rate and length; the Python runner reads wav and flac at any sample rate. The first run prepares the package for the Neural Engine once, which can take a minute or two; the compiled copy is kept, so later loads take under a second.
Phonon-2 also runs through the fermion command line on Apple silicon, Linux, Windows and NVIDIA GPUs. See the
Phonon-2 card.
Long audio
An utterance of up to 35 seconds is read in one pass by the smallest encoder window that holds it. Longer recordings are cut at pauses into windows of up to 15 seconds and their words are joined, so no word is split at a boundary. Pass recordings whole; both runners do the cutting.
Accuracy
Every utterance of the seven English test sets, scored with the Open ASR Leaderboard's own code, with the Core ML package beside the Phonon-2 MLX engine.
| Set | Core ML, Neural Engine | Phonon-2 MLX engine |
|---|---|---|
| LibriSpeech clean | 1.73 | 1.72 |
| LibriSpeech other | 3.91 | 3.92 |
| AMI | 9.33 | 9.37 |
| Earnings-22 | 6.99 | 6.96 |
| GigaSpeech | 8.33 | 8.35 |
| SPGISpeech | 3.70 | 3.70 |
| VoxPopuli | 2.46 | 2.46 |
| Mean of seven | 5.21 | 5.21 |
Word error rate, %, lower is better. Phonon-2 transcribes English. The Phonon-2 card lists its accuracy in other languages.
Speed
On an M5 MacBook Air, with the model loaded once, the encoder on the Neural Engine and the decoder on the CPU.
| Audio | Time | Times real time |
|---|---|---|
| One hour of LibriSpeech read as a single file | 6.0 s | 606× |
| 158 hours of test-set audio, utterance by utterance | 31 min | 305× |
| 20 recordings of 2 seconds to 2 minutes (797 s) | 2.0 s | 400× |
| A recording of 3 to 5 seconds, model loaded (median of 20) | 11 ms |
Files
| File | Size | Contents |
|---|---|---|
Phonon-2.mlpackage |
332 MB | the encoder for 5, 10, 15 and 35 second windows over one shared weight set (five-value weights stored exactly, 16-bit float activations); audio is masked to its true length inside the window |
decoder.bin |
13 MB | the prediction network, joint and vocabulary |
manifest.json |
source checksum, window lengths and input contract | |
phonon_coreml.py |
the Python runner, with frontend.py and segmenter_np.py |
|
phonon-coreml-src.tgz |
the Swift package and command-line tool | |
ci/ |
three LibriSpeech clips and their transcripts, for a quick check | |
SHA256SUMS |
a checksum for every file above |
Licence
The weights are released under CC-BY-4.0. They derive from NVIDIA's parakeet-tdt-0.6b-v3, whose tokenizer and output
conventions (punctuation, casing, numerals) they keep, and NOTICE lists the changes. The Swift package, the Python runner and
this repository's code are released under Apache 2.0.
- Downloads last month
- 37