Nemotron 3 Diarization — Core AI
FP16 Apple Core AI conversion of NVIDIA Nemotron 3 Diarization,
source revision f667ed73aee57d40cc39428eb768b4fd87a0a29e, under
OpenMDW-1.1. See LICENSE and NOTICE.
This conversion is not endorsed by NVIDIA.
Two self-contained bundles preserve streaming speaker identities and support up to eight overlapping speakers:
| Bundle | New audio per step | Lookahead | Encoder capacity | Speaker cache / FIFO |
|---|---|---|---|---|
fast32/, live |
2.56 s | 0.32 s | 352 positions | 264 / 40 positions |
fast128/, final |
10.24 s | 0.32 s | 448 positions | 264 / 40 positions |
An encoder position represents 80 ms of audio. Both policies update their speaker
cache every 40 positions. Initial buffering is approximately 2.88 / 10.56 seconds.
These are explicit profiles, not the upstream processor's default low-latency
or offline settings. Each bundle is about 199.4 MB and contains source .aimodel
assets; there are no architecture-specific ahead-of-time compiled assets.
Requires a physical Apple-silicon device with macOS 27 or iOS 27, Core AI, and a
model-specific host runtime. These graphs consume features/state, not audio files.
The runtime must implement the source FFT/log-mel frontend, feature stacking,
arrival-order speaker cache, FIFO, lookahead and final flush. metadata.json
records the versioned contract; mel-filter.f32, window.f32 and silence.f32
provide frontend weights and the learned silence embedding. Geometry must be read
from this metadata rather than inferred from filenames.
embedding.aimodel maps stacked features [1,1024,1,chunk_frames+4] to hidden
embeddings [1,512,1,chunk_frames+4]. encoder.aimodel receives hidden input
embeddings [1,512,1,encoder_capacity] and a valid-prefix mask
[1,1,1,encoder_capacity]; it returns logits [1,8,8,encoder_capacity].
The last two axes are subframes and encoder positions. Convert to time-major
10 ms frames, apply independent sigmoid probabilities, remove cached positions
and defer lookahead until it becomes committed audio. Preserve the source
cache's selection and ordering. The graphs preserve full bidirectional attention;
they do not truncate context or discard input audio.
Measured Mac performance
M3 MacBook Air, 16 GB, macOS 27 build 26A428. Three-run medians after warmup, with nominal thermals, delivering preloaded mono 16 kHz PCM in 2.56-second blocks without real-time pauses. Frontend, graph execution, speaker-cache updates and probability delivery are included; preparation and file I/O are excluded. The input is JFK's “We choose to go to the Moon” speech: a 20-second excerpt and the 18-minute 15-second recording. RTFx is audio duration divided by processing time; higher is faster.
| Policy | Input | Core AI | Core AI RTFx | FluidAudio | Speedup |
|---|---|---|---|---|---|
| fast32 | 20 seconds | 0.101 s | 197.5× | 0.131 s | 1.30× |
| fast32 | 18 min 15 s | 5.445 s | 201.2× | 7.430 s | 1.36× |
| fast128 | 20 seconds | 0.038 s | 531.9× | 0.062 s | 1.65× |
| fast128 | 18 min 15 s | 2.103 s | 520.9× | 3.551 s | 1.69× |
The FluidAudio baseline uses its public API unchanged, monolithic fast32/fast128
models, and all Core ML compute units. Dependency revision:
5c51c5c93afff0d89594a2a93c3103e790ba648c; model repository:
FluidInference/nemotron-3-diarization-coreml, revision
25a90f97f254428d4b30374b76af9c74fdee8327. Native/source output contains 109,531
valid frames on the full recording; FluidAudio emits two extra tail frames.
These processing rates do not measure energy use or live algorithmic latency.
Measured iPhone performance
iPhone 15 Pro Max, iOS 27 build 24A437, Release runtime. Three-run medians after warmup, nominal thermals, no real-time pauses. These runs include bounded WAV reading/conversion, frontend, graph execution, speaker-cache updates and output delivery. Preparation, warmup and writing result files are excluded. The same 20-second excerpt and full “We choose to go to the Moon” speech are used.
| Policy | 20 seconds | RTFx | 18 min 15 s | RTFx |
|---|---|---|---|---|
| fast32 | 0.104 s | 192.7× | 5.532 s | 198.0× |
| fast128 | 0.041 s | 490.0× | 2.129 s | 514.4× |
First observed complete preparations took 24.45 s / 26.20 s (fast32 / fast128); subsequent preparations were approximately 0.04–0.16 s. These measurements include any requested device specialization and loading, and do not guarantee a pristine-cache compilation time on other installations.
Separate full-recording traces contain all 428 / 107 update scopes, with ANE predictions inside every update and zero target-process GPU intervals. This demonstrates ANE execution throughout both policies, not ALU occupancy or the absence of host CPU work. The recorded peak client footprint was about 290 / 285 MB, including allocations charged to the app; unattributed driver and system memory are not included. Clean timings above exclude Instruments.
Both short and full phone probability arrays are bit-identical to the qualified Mac outputs. The source-agreement limits below therefore also apply to these phone results. No phone energy measurement or phone FluidAudio comparison is claimed.
Device specialization and placement
First observed preparations of the equivalent FP16 authoring graphs took 23.99 s / 39.67 s (fast32 / fast128). The distributed assets, after removal of authoring debug information, loaded in 3.71 s / 3.69 s with related driver caches already present. Cached launches typically take 0.02–0.08 s. These observations are not controlled clean-install compilation measurements; cache reuse prevents claiming the shorter load as a cold-start result.
Separate bounded hardware traces on a 60-second prefix show all 24 fast32 and six fast128 updates containing the expected two ANE predictions each, with no GPU intervals for the target process. This establishes ANE placement, not ALU occupancy. Initial full-file endpoint client-plus-attributed-neural memory was approximately 442 / 436 MB, including preloaded PCM; it is not peak system memory.
Numerical checks
The independent oracle is the original FP32 checkpoint through Transformers
revision 07338b6c74a578868368e6e549dea83414e4b8cb. The FP32 authoring graph agrees
within 0.000001 relative RMS, and native FP16 frozen logit fixtures are within
0.06–0.21% relative RMS. Full streaming checks include cache growth/compression,
frontend behavior and final flush. Changing PCM delivery between 80 ms,
2.56 seconds and an entire 20-second clip produces identical native probabilities.
The packaged graphs produce bit-identical output to the qualified authoring graphs.
Thresholded speaker/frame disagreement versus FP32, relative to the union of active decisions at threshold 0.5:
| Policy | Full Moon speech | 97.6s multi-speaker fixture |
|---|---|---|
| fast32 | 1.32% | 0.085% |
| fast128 | 0.198% | 0.105% |
The multi-speaker fixture is diarization_example.mp3 from
hf-internal-testing/dummy-audio-samples. Both policies preserve their source's
active speaker-slot counts there. Full-recording fast32 still has 36 threshold
crossings in a second slot where FP32 has none. These are source-agreement
measurements, not human-annotated diarization error rates. FP16 is selected
because INT8 introduced a substantial fast32 speaker-slot error without a useful
throughput gain. Quality depends on the source model and the selected policy.
Model tree for coder543/nemotron-3-diarization-coreai
Base model
nvidia/Nemotron-3-Diarization