Nemotron English โ Core AI
Public Apple Core AI assets derived from nvidia/nemotron-speech-streaming-en-0.6b, revision ebe59e5a817142986528bbbee5dba8db7b38ed50.
Targets physical Apple silicon devices running macOS 27 or iOS 27. These are source .aimodel assets that specialize on the device.
The bundle shares weights across 1.12-second quality and 160-millisecond preview entrypoints. It contains a W8A16 streaming encoder, FP16 causal subsampling, FP16 recurrent predictor and vocabulary graphs, embeddings, vocabulary and frontend coefficients. Input is continuous 16 kHz mono audio. Attention, convolution, subsampling and recurrent decoder state persist across updates. metadata.json specifies runtime geometry, the default quality entrypoints and optional preview entrypoints.
These are model assets. A host must implement continuous feature extraction, preserve graph states, and perform greedy recurrent decoding. Quality-only hosts can use the ordinary entrypoints without running the preview branch.
Optional low-latency preview
The measured dual-mode host maintains independent quality and preview state. It runs six 160 ms speculative updates between 1.12 s quality updates, then replaces speculative state with a copy of the quality state. Only words beyond the quality transcript are displayed provisionally; quality updates replace that suffix, and finalization returns the quality transcript. Speculative state never feeds the quality decoder. Catch-up can skip preview work entirely. A display policy can withhold speculative sentence punctuation until quality confirms it to avoid flickering punctuation.
This trades additional work for earlier provisional text. It is optional; short updates are memory-heavy and are not seven times cheaper than quality updates. Update intervals are not measured word latency, and these benchmarks do not establish energy consumption.
Measured performance
John F. Kennedy's โWe choose to go to the Moonโ speech: a 20-second excerpt and the full 18 minutes 15 seconds (1,095.3 s). Release host runtimes; no real-time pauses. Results include frontend, persistent graph states, recurrent decoding, partial delivery and final flush. Preparation and warmup are excluded. Short results are medians of three repetitions; full results are one measured repetition after warmup. Higher RTFx is faster.
M3 MacBook Air, 16 GB, macOS 27.0.1 (26A434). PCM loading is excluded; the full recording had a full-recording warmup.
| Mode | 20 s excerpt | Full speech | Full RTFx |
|---|---|---|---|
| Quality 1.12 s | 0.277 s | 17.601 s | 62.2ร |
| Standalone 160 ms | 1.516 s | 89.992 s | 12.2ร |
| Dual 160 ms + 1.12 s | 1.588 s | 90.922 s | 12.0ร |
iPhone 15 Pro Max (A17 Pro), iOS 27.0.1. Includes bounded WAV reading/conversion; 20-second warmup. All measured passes remained foregrounded at nominal thermal state.
| Mode | 20 s excerpt | Full speech | Full RTFx |
|---|---|---|---|
| Quality 1.12 s | 0.359 s | 21.070 s | 52.0ร |
| Dual 160 ms + 1.12 s | 1.990 s | 117.545 s | 9.3ร |
Full dual processing performs 979 quality updates and 5,868 preview updates. Peak measured dual-mode client footprint on iPhone was 756.6 MB, excluding allocations in accelerator services.
For an independent speed reference, an earlier M3 run of FluidAudio's 1.12 s tier measured 0.374 s short and 22.479 s full (48.7ร), three-run medians with preloaded PCM. It was not rerun alongside the dual-bundle measurements above. FluidAudio library revision: 5c51c5c93afff0d89594a2a93c3103e790ba648c; converted baseline revision: e673531caa6d25ab7baf5a8c14c9b99ba1551838. Its source checkpoint identity was not independently verified. The official NVIDIA checkpoint is the correctness reference.
Specialization and cached loading
| Device | Initial preparation | Fresh-process cached preparation |
|---|---|---|
| M3 MacBook Air | 51.5 s median; 51.4โ52.1 s | 31.8 ms median |
| iPhone 15 Pro Max | 55.3 s observed | 138โ190 ms |
The Mac measurements used three trials with the public model-specific Core AI cache deletion API, then loaded every exported function. Encoder preparation accounted for 50.5 s of the median. The phone encoder accounted for 53.1 s of its initial preparation. Underlying compiler-cache reuse is not fully observable: these are not guaranteed pristine-install times. The extra preview shape increases specialization work even when a host chooses quality-only inference.
Placement and correctness
A separate Mac short trace captured 523 ANE predictions across 523 graph calls, including 108 preview encoders and 18 quality encoders, with zero target-process GPU intervals. This trace preceded a host-only state-copy optimization; its wall time is not the final timing above.
The iPhone short trace captured ANE predictions for all 108 preview encoders, 18 quality encoders, 126 subsampling calls and 147 predictor steps, with zero target-process GPU intervals. Only 1 of 124 tiny joint-only scopes contained an ANE prediction, so the trace does not establish the device for every joint-only invocation. Public hardware events are attributed by temporal containment; these counts do not measure ALU occupancy, memory bandwidth or energy use.
Every authoritative partial token/frame history and final token/frame sequence in dual mode matches quality-only within this bundle on both tested recordings, on Mac and iPhone. Mac checks additionally covered paced delivery and catch-up that bypasses preview.
Quality transcript text and token IDs match the separately qualified single-shape conversion. Re-exporting changed 14 of 4,413 full-recording Mac token emission frames, mostly by one 80 ms frame and at most four; one partial token history differed before converging. Token emission times are not forced-aligned word boundaries.
Official-source qualification used Transformers 5.16.0 FP32 CPU/SDPA streaming generation from the NVIDIA weights. The matching full-speech quality transcript scores 2.48% (55/2,220) Whisper-normalized WER, versus 2.52% (56/2,220) for that FP32 source. The short quality excerpt matches official-source token IDs. Reference transcript. One recording does not establish a general accuracy ranking; provisional 160 ms output has not been independently qualified for WER.
License
Distributed under nvidia-open-model-license; see LICENSE and NOTICE. Hugging Face provides file checksums for each repository snapshot.
Model tree for coder543/nemotron-speech-streaming-en-0.6b-coreai
Base model
nvidia/nemotron-speech-streaming-en-0.6b