Canary-1B-v2 Speech Translation — CoreML (ANE)

On-device speech-to-text translation (English ↔ 24 European languages) on the Apple Neural Engine, using the exact CoreML models published at FluidInference/canary-1b-v2-coreml.

No extra weights are needed: nvidia/canary-1b-v2 is a multi-task attention encoder-decoder whose task is selected by two decoder prompt slots. Transcription is the special case source == target; setting a different target language token makes the same models translate the speech. This repo documents that usage and hosts the benchmark artifacts; the model files live in (and are auto-downloaded from) the main repo.

Prompt contract

▁ <|startofcontext|> <|startoftranscript|> <|emo:undefined|> <SOURCE> <TARGET> <|pnc|> <|noitn|> <|notimestamp|> <|nodiarize|>

English→German is one token away from English transcription:

transcribe en:  [16053, 7, 4, 16, 64, 64, 5, 9, 11, 13]
translate en→de:[16053, 7, 4, 16, 64, 78, 5, 9, 11, 13]

prompt_tokens.json in this repo maps all 25 language codes to their token ids.

Usage (FluidAudio)

Swift (FluidAudio, branch with Canary translation support):

let manager = try await CanaryManager.load(source: .english, target: .german)
let german = try await manager.transcribe(audioURL: url)  // ≤ 15 s window

CLI:

swift run fluidaudiocli canary-transcribe audio.wav --translate-to de --source-lang en

Benchmarks (FLEURS test, int4/ANE, Apple M5 Pro)

Scored with sacreBLEU / chrF2 / COMET (Unbabel/wmt22-comet-da). References are the raw cased FLEURS transcriptions of the aligned target-language sentences.

Direction Pairs BLEU chrF2 COMET RTFx
en→de 641 29.98 59.86 82.70 10.3×
de→en 862 33.06 62.15 83.50 14.7×

For reference, NVIDIA reports FLEURS 24-language fp16 averages of BLEU 29.4 / COMET 84.56 (en→X) and BLEU 29.08 / COMET 79.30 (X→en) for the base model.

The 15-second window matters

The CoreML contract is a fixed 15 s window; longer audio is chunked with an overlap-merge designed for transcripts, which is unreliable for translation (translated text is not token-stable across window seams). Splitting the benchmark by clip length:

Subset en→de COMET de→en COMET
≤ 15 s (single window) 83.44 85.67
> 15 s (chunked) 73.88 78.59

Single-window quality matches the published fp16 numbers — the int4 quantization penalty is ≤ ~1 COMET point. For long audio, segment at pauses into ≤ 15 s pieces (e.g. with a VAD) instead of relying on the chunked path.

fleurs_{en_de,de_en}_hypotheses.json contain the parallel hypotheses and references used for these scores.

Files

File Contents
prompt_tokens.json prompt layout + language-code → token-id map + example
fleurs_en_de_hypotheses.json 641 en→de hypotheses/references (FLEURS test)
fleurs_de_en_hypotheses.json 862 de→en hypotheses/references (FLEURS test)

Attribution

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for FluidInference/canary-speech-translation-coreml

Quantized
(11)
this model