Parakeet v3 โ€” Apple Core AI streaming conversion

Conversion of nvidia/parakeet-tdt-0.6b-v3 at 541d1f99c6b0c3cd0b11a95167540bb8edefd82b for Apple's CoreAISpeech runtime on macOS 27. Source training is by NVIDIA (v3) and Moondream (Ultra/Redux). Conversion by maxffarrell; no training or endorsement.

Weights are CC-BY-4.0. Preserve attribution to NVIDIA and, for Ultra/Redux, Moondream when redistributing. License terms. The export recipe is Apple BSD-3-Clause; see APPLE_LICENSE.

Format and use

Download model.zip, verify against artifact.json, and extract. Pass the extracted parakeet-tdt-0.6b-v3_float16_streaming150 directory to SpeechRecognitionModel(resourcesAt:) from apple/coreai-models at 52c84ba. Three .aimodel graphs: encoder, recurrent decoder step, joint. Use startStream, append(pcm:) and finishStream; collect all finalized update segments, since finishStream returns only its last segment.

FP16 static window: 128 mel bins ร— 1201 frames; left/chunk/right encoder frames 113/12/25. Audio: mono Float32 16 kHz. Buffered streaming has 2.96 seconds of theoretical audio latency before inference or display stabilization. Streaming is buffered full-window inference, not a cache-aware causal model. No guaranteed exclusive Neural Engine placement or speedup.

Changes and caveats

Graph format, FP16 precision and fixed streaming geometry differ from the source. Redux's ternary base-3 packing is expanded exactly before FP16 conversion: the original 178 MB packed size and Photon specialized performance are not retained. All conversions are about 1.2 GB unpacked. Auxiliary Photon VAD tensors are omitted; CoreAISpeech endpointing is used separately. NVIDIA v3 tokenizer vocabulary was checked against each source.

Validation

conversion-validation.json compares all three graphs against the pinned PyTorch source, including masked encoder frames, nonzero recurrent state and joint logits, using the macOS Core AI GPU preference. Acceptance: cosine > 0.995 and relative RMSE < 0.05. Swift runtime transcription and stream boundary regression checks are recorded in the Pladder fork. These are conversion checks on synthetic speech, not multilingual quality claims or calibrated confidence scores.

Reproduce

See tools/coreai/export_models.py, verify_models.py and the dependency pins in Pladder. Source and recipe revisions, omitted tensors and precision are embedded in provenance.json. Every bundle file has a SHA-256 entry in SHA256SUMS.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support

Model tree for maxffarrell/parakeet-v3-coreai

Finetuned
(106)
this model