Parakeet TDT 0.6B V3 โ Core AI
Apple Core AI conversion of NVIDIA Parakeet V3,
pinned to 541d1f99c6b0c3cd0b11a95167540bb8edefd82b. Converted weights retain the upstream
Creative Commons Attribution 4.0 license.
See LICENSE and NOTICE. This conversion is not endorsed by NVIDIA.
Choose one self-contained bundle:
fast/: full finite Fourier relative attention, 192 encoder positions, fixed 15-second chunks, and an FP16 ANE decoder batching up to 128 independent chunks.quality/: full sinusoidal relative attention, shared weights for 576 and 3,008 encoder positions (45-second and four-minute inputs), exact tiled convolution subsampling, fixed 240-second chunks, and an FP32 CPU decoder batching up to four independent chunks. Select the smallest shape that fits each chunk; short recordings use the 45-second shape.
Both use W8A16 encoder weights with selected sensitive projections retained in
FP16, and four concurrent encoder requests. Neither truncates attention within
its chunk or drops input audio. metadata.json supplies available input shapes;
runtime.json records the qualified scheduling policy. Weights are shared across
static functions in each asset. The labels describe context and execution policy,
not a universal quality ranking: longer context did not improve WER in the English
recording measured below.
Requires physical Apple silicon on macOS 27 or iOS 27 and a model-specific host runtime. Graphs accept model features and recurrent states, not audio files. The host must implement the frontend, greedy TDT loop and tokenizer described by the metadata and sidecars. Vocabulary: 8,192 nonblank tokens, blank ID 8,192, durations 0/1/2/3/4, and at most ten symbols per frame. The FP16 decoder returns partition/local token-ID pairs so IDs above 2,048 are reconstructed without rounding. Recurrent state is independent between chunks.
Quality host contract: subsampling_output_scale is 16. Multiply the
subsampling graph's output by this factor on the host before passing it into
the encoder. This power-of-two scaling protects the FP16 projection from overflow;
feeding the scaled output directly into the encoder produces incorrect results.
Both graph functions use this contract. Fast defaults to scale 1.
TDT emission frames and predicted durations support token/word timings at an 80 ms frame step. These are native alignments, not forced alignment.
Measured performance
M3 MacBook Air (16 GB), macOS 27 build 26A428. Medians of three runs after warmup; frontend, encoder, host rescaling and decoding are included. Preparation, file I/O and chunk planning are excluded. Audio is JFK's โWe choose to go to the Moonโ speech: a 20-second excerpt and the 18-minute 15-second recording.
| Bundle | Audio duration | Transcription | Audio / elapsed |
|---|---|---|---|
| fast | 20 seconds | 0.129 s | 155.5ร |
| fast | 18 min 15 s | 1.942 s | 563.9ร |
| quality | 20 seconds | 0.189 s | 106.0ร |
| quality | 18 min 15 s | 7.973 s | 137.4ร |
Tokens and native word timings were stable across runs, with nominal thermal state. The full recording uses 74 fast chunks or five quality chunks. These contexts and execution policies differ.
Whisper-normalized WER against the supplied long-recording reference:
| Bundle / reference | Word errors | WER |
|---|---|---|
| fast | 79/2220 | 3.56% |
| FP32 source, same 15-second cuts | 77/2220 | 3.47% |
| quality | 81/2220 | 3.65% |
| FP32 source, same four-minute cuts | 96/2220 | 4.32% |
With the benchmark's simpler normalization, fast scores 95/2219 and its source 94/2219; quality scores 91/2219 and its source 107/2219. This is one English recording, not a general accuracy ranking. The reference transcript is from stt-bench-matrix. The numerical oracle uses the independently loaded upstream FP32 weights.
Quality's two shapes produce identical valid outputs on the 20-second numerical fixture, with 2.52% encoder relative RMS error versus FP32. Fast's 15.35-second fixture measures 2.99%. All frontend, subsampling and encoder intermediates were finite in separate full-recording checks. Both bundles exactly reproduce the source tokens on additional German and Ukrainian samples; their encoder errors range from 2.33% to 3.60%. These checks do not establish accuracy across every supported language. Quantization and floating-point operation order can change token decisions, including a proper name in the quality English fixture.
Word timings are ordered, bounded and text-preserving; human word-boundary accuracy was not measured. Decoder tests cover all 8,193 token IDs and recurrent FP16-versus-FP32 steps, including odd IDs above the FP16 exact-integer range.
Bounded full-recording traces contain ANE predictions within every encoder and subsampling call: 148 for fast and ten for quality. Fast's batched decoder loop contains 94 further ANE predictions. Neither trace contains target-process GPU intervals. Quality decoding uses the CPU. Compiler manifests mark the used encoder/subsampling and batched decoder functions fully placed on ANE. This is placement evidence, not arithmetic utilization. Quality's sampled peak client-plus-attributed-neural memory is about 1.29 GB, excluding unattributed compiler, driver and system memory.
Comparison with Apple's Core AI Parakeet V3 example
Measured September 28, 2026 on the same M3 MacBook Air (16 GB, macOS 27 build
26A428), using Apple's unmodified
coreai-models implementation at e7b24da.
Both conversions use the same NVIDIA source revision listed above.
At matched 15-second chunk boundaries, this repository's Fast bundle completed the full Moon speech 6.5ร faster, with identical Whisper-normalized WER. The five-second Apple default scored worse on this recording; that is a comparison of default configurations, not evidence of a generally more accurate underlying model.
| Implementation / configuration | 20-second excerpt | Complete speech | Complete-speech RTFx | Complete-speech WER |
|---|---|---|---|---|
| Apple default: FP32, 5-second shape | 0.342 s | 18.653 s | 58.7ร | 11.13% |
| Apple: FP16, 15-second shape | 0.336 s | 13.361 s | 82.0ร | 3.56% |
| This repository: Fast, W8A16, 15-second shape | 0.243 s | 2.048 s | 534.8ร | 3.56% |
| This repository: Quality, W8A16, 45/240-second shapes | 0.187 s | 7.852 s | 139.5ร | 3.65% |
These are fresh comparison runs, distinct from the qualification measurements above. Each entry is the median of three complete passes after warmup. Timing includes frontend, encoder, decoding and transcript construction; it excludes preparation, file decoding and chunk planning. All implementations process the complete 1,095.32-second recording. Fast uses 74 chunks; Quality uses five. Apple's five-second configuration uses 220 chunks. Apple's runs were at nominal thermal state; Fast's full-recording runs transitioned from nominal to fair. The short-clip Fast result is slower than its earlier qualification measurement; these measured results are reported without substituting the earlier best run.
Apple's static offline API pads or truncates PCM to its exported window. The
comparison caller therefore divides the entire recording into consecutive,
non-overlapping windows and invokes the original API sequentially, resetting
its decoder between windows. Without that caller-side chunking, passing the
complete recording to the five-second static export would drop nearly all its
audio. Apple was exported with the original script and its declared dependencies:
default flags for FP32/5s, or --dtype float16 --audio-seconds 15 for FP16/15s.
Both Swift hosts were built in Release configuration.
WER uses the same reference linked above and Whisper English normalization: Apple 5s = 247/2,220 errors; Apple 15s = 79/2,220; Fast = 79/2,220; Quality = 81/2,220. With the benchmark's simpler normalization the corresponding counts are 288/2,219, 96/2,219, 95/2,219 and 91/2,219. The one-error advantage under that normalization is too small to establish an accuracy improvement. All three repetitions of each configuration returned identical text. This is one English recording; it does not establish a multilingual accuracy ranking. Changing Apple's default to the tested variant changes both precision and chunk length, so their individual effects are not isolated.
Runtime and feature tradeoffs
| Area | These bundles and qualified host runtime | Apple's tested example |
|---|---|---|
| Encoder storage | W8A16 with sensitive projections retained in FP16 | FP32 by default; FP16 optional |
| Offline scheduling | Four concurrent encoders; Fast batches the ANE decoder across up to 128 independent chunks | Sequential caller-side chunks; scalar predictor and joint calls |
| Long recordings | Bounded chunk processing that retains all input audio | Static offline call is window-limited; caller must split longer audio |
| Word timings | Native TDT token/word timings with global chunk offsets | Public offline API returns text and decode statistics, without word timings |
| Context choices | Fast 15s; Quality shares weights across 45s and 240s functions | One configurable static duration, or a symbolic-length encoder export |
| Streaming | These bundles are qualified for final transcription | Separate buffered-window streaming export and runtime are available |
The host features require a compatible runtime; downloading .aimodel files
alone does not implement chunking or word timing. The matched-length speedup
includes graph design, quantization and scheduling differences, not just an
accelerator comparison. Apple's loader uses Core AI's default device selection;
its actual placement was not traced in this experiment. The ANE placement
qualification for these bundles is documented above.
Apple's dynamic encoder may avoid fixed-window padding, but its README warns that the FP32 dynamic GPU path is unreliable at many shapes. Dynamic and streaming exports were inspected, not benchmarked here. Our Quality configuration offers longer offline context at a throughput and specialization cost; its name does not imply better WER on every recording.
Measured iPhone performance
iPhone 15 Pro Max, iOS 27 build 24A437, Release runtime. Three-run medians after warmup, using the same Moon speech. Includes bounded WAV reading/conversion, chunk planning, frontend, encoder, decoding and output delivery. Preparation, warmup and report writing are excluded. Hardware traces were captured separately.
| Bundle | Audio | Transcription | Audio / elapsed |
|---|---|---|---|
| fast | 20 seconds | 0.162 s | 123.1ร |
| fast | 18 min 15 s | 2.498 s | 438.5ร |
| quality | 20 seconds | 0.212 s | 94.4ร |
| quality | 18 min 15 s | 8.767 s | 124.9ร |
| Bundle | Observed preparation with specialization | Cached preparation |
|---|---|---|
| fast | 46.6 s | 0.622 s |
| quality | 489.9 s | 0.089 s |
Preparation includes loading and any specialization requested by Core AI. Underlying driver cache reuse is opaque; these are observed histories, not controlled fresh-install measurements or evidence of an empty system cache. The quality inference runs follow a cooldown after compilation. All reported passes remained foregrounded; nominal and fair thermal states were observed.
Fast's complete recording trace contains 242 unique ANE predictions across 74 encoder calls, 74 subsampling calls and the batched decoder loop, with zero target-process GPU intervals. Fast token IDs and transcript text match the qualified Mac outputs for both input lengths.
Quality's encoder and subsampling calls contain ANE predictions for both shapes, with zero target-process GPU intervals. Quality decoding uses the CPU.
Quality's short output matches the Mac. Full-recording tokens differ but are stable across all phone runs: 65/2,220 errors (2.93% Whisper-normalized WER), versus 81/2,220 (3.65%) on the Mac and 96/2,220 (4.32%) for the matching FP32 source. With the benchmark's simpler normalization, the phone scores 76/2,219 and the Mac 91/2,219. Phone testing here covers this English recording; the additional German and Ukrainian source checks above were run on Mac. This single-recording difference does not establish a general device or quality ranking.
Placement traces establish execution, not ALU utilization or energy use. Concurrent scopes can overlap the same hardware event; aggregate counts use unique intervals.
Device specialization and loading
Source .aimodel files specialize on the device. On the same Mac:
| Bundle | Observed preparation with specialization | Subsequent cached preparation |
|---|---|---|
| fast | 34 s | 0.10โ0.21 s |
| quality | encoder 5 min 32 s; subsampling 32 s | 0.02โ0.05 s |
The quality components were measured in separate processes: their specialization times sum to about 6 min 4 s, not a measured single cold whole-bundle load. The encoder measurement uses the exact distributed graph hash. Fast's observed 34-second preparation also has prior cache history. Neither is a controlled fresh-install measurement. Active OS caches were preserved.
Preparation includes Core AI model initialization and function loading; it excludes host sidecar reads, audio I/O, chunk planning, warmup and transcription. These are observed cache histories, not guaranteed device load times. The underlying ANE cache state is not fully observable, and OS/device/application changes can require specialization again.
AoT compiled artifacts are omitted because no significant load-time benefit has been demonstrated. Authoring debug locations were removed while preserving graph signatures and operation counts. Hugging Face provides file checksums for each repository snapshot.
Model tree for coder543/parakeet-v3-coreai
Base model
nvidia/parakeet-tdt-0.6b-v3