Nemotron 3 Diarization — Core AI

FP16 Apple Core AI conversion of NVIDIA Nemotron 3 Diarization, source revision f667ed73aee57d40cc39428eb768b4fd87a0a29e, under OpenMDW-1.1. See LICENSE and NOTICE. This conversion is not endorsed by NVIDIA.

Two self-contained bundles preserve streaming speaker identities and support up to eight overlapping speakers:

Bundle New audio per step Lookahead Encoder capacity Speaker cache / FIFO
fast32/, live 2.56 s 0.32 s 352 positions 264 / 40 positions
fast128/, final 10.24 s 0.32 s 448 positions 264 / 40 positions

An encoder position represents 80 ms of audio. Both policies update their speaker cache every 40 positions. Initial buffering is approximately 2.88 / 10.56 seconds. These are explicit profiles, not the upstream processor's default low-latency or offline settings. Each bundle is about 199.4 MB and contains source .aimodel assets; there are no architecture-specific ahead-of-time compiled assets.

Requires a physical Apple-silicon device with macOS 27 or iOS 27, Core AI, and a model-specific host runtime. These graphs consume features/state, not audio files. The runtime must implement the source FFT/log-mel frontend, feature stacking, arrival-order speaker cache, FIFO, lookahead and final flush. metadata.json records the versioned contract; mel-filter.f32, window.f32 and silence.f32 provide frontend weights and the learned silence embedding. Geometry must be read from this metadata rather than inferred from filenames.

embedding.aimodel maps stacked features [1,1024,1,chunk_frames+4] to hidden embeddings [1,512,1,chunk_frames+4]. encoder.aimodel receives hidden input embeddings [1,512,1,encoder_capacity] and a valid-prefix mask [1,1,1,encoder_capacity]; it returns logits [1,8,8,encoder_capacity]. The last two axes are subframes and encoder positions. Convert to time-major 10 ms frames, apply independent sigmoid probabilities, remove cached positions and defer lookahead until it becomes committed audio. Preserve the source cache's selection and ordering. The graphs preserve full bidirectional attention; they do not truncate context or discard input audio.

Measured Mac performance

M3 MacBook Air, 16 GB, macOS 27 build 26A428. Three-run medians after warmup, with nominal thermals, delivering preloaded mono 16 kHz PCM in 2.56-second blocks without real-time pauses. Frontend, graph execution, speaker-cache updates and probability delivery are included; preparation and file I/O are excluded. The input is JFK's “We choose to go to the Moon” speech: a 20-second excerpt and the 18-minute 15-second recording. RTFx is audio duration divided by processing time; higher is faster.

Policy Input Core AI Core AI RTFx FluidAudio Speedup
fast32 20 seconds 0.101 s 197.5× 0.131 s 1.30×
fast32 18 min 15 s 5.445 s 201.2× 7.430 s 1.36×
fast128 20 seconds 0.038 s 531.9× 0.062 s 1.65×
fast128 18 min 15 s 2.103 s 520.9× 3.551 s 1.69×

The FluidAudio baseline uses its public API unchanged, monolithic fast32/fast128 models, and all Core ML compute units. Dependency revision: 5c51c5c93afff0d89594a2a93c3103e790ba648c; model repository: FluidInference/nemotron-3-diarization-coreml, revision 25a90f97f254428d4b30374b76af9c74fdee8327. Native/source output contains 109,531 valid frames on the full recording; FluidAudio emits two extra tail frames. These processing rates do not measure energy use or live algorithmic latency.

Measured iPhone performance

iPhone 15 Pro Max, iOS 27 build 24A437, Release runtime. Three-run medians after warmup, nominal thermals, no real-time pauses. These runs include bounded WAV reading/conversion, frontend, graph execution, speaker-cache updates and output delivery. Preparation, warmup and writing result files are excluded. The same 20-second excerpt and full “We choose to go to the Moon” speech are used.

Policy 20 seconds RTFx 18 min 15 s RTFx
fast32 0.104 s 192.7× 5.532 s 198.0×
fast128 0.041 s 490.0× 2.129 s 514.4×

First observed complete preparations took 24.45 s / 26.20 s (fast32 / fast128); subsequent preparations were approximately 0.04–0.16 s. These measurements include any requested device specialization and loading, and do not guarantee a pristine-cache compilation time on other installations.

Separate full-recording traces contain all 428 / 107 update scopes, with ANE predictions inside every update and zero target-process GPU intervals. This demonstrates ANE execution throughout both policies, not ALU occupancy or the absence of host CPU work. The recorded peak client footprint was about 290 / 285 MB, including allocations charged to the app; unattributed driver and system memory are not included. Clean timings above exclude Instruments.

Both short and full phone probability arrays are bit-identical to the qualified Mac outputs. The source-agreement limits below therefore also apply to these phone results. No phone energy measurement or phone FluidAudio comparison is claimed.

Device specialization and placement

First observed preparations of the equivalent FP16 authoring graphs took 23.99 s / 39.67 s (fast32 / fast128). The distributed assets, after removal of authoring debug information, loaded in 3.71 s / 3.69 s with related driver caches already present. Cached launches typically take 0.02–0.08 s. These observations are not controlled clean-install compilation measurements; cache reuse prevents claiming the shorter load as a cold-start result.

Separate bounded hardware traces on a 60-second prefix show all 24 fast32 and six fast128 updates containing the expected two ANE predictions each, with no GPU intervals for the target process. This establishes ANE placement, not ALU occupancy. Initial full-file endpoint client-plus-attributed-neural memory was approximately 442 / 436 MB, including preloaded PCM; it is not peak system memory.

Numerical checks

The independent oracle is the original FP32 checkpoint through Transformers revision 07338b6c74a578868368e6e549dea83414e4b8cb. The FP32 authoring graph agrees within 0.000001 relative RMS, and native FP16 frozen logit fixtures are within 0.06–0.21% relative RMS. Full streaming checks include cache growth/compression, frontend behavior and final flush. Changing PCM delivery between 80 ms, 2.56 seconds and an entire 20-second clip produces identical native probabilities. The packaged graphs produce bit-identical output to the qualified authoring graphs.

Thresholded speaker/frame disagreement versus FP32, relative to the union of active decisions at threshold 0.5:

Policy Full Moon speech 97.6s multi-speaker fixture
fast32 1.32% 0.085%
fast128 0.198% 0.105%

The multi-speaker fixture is diarization_example.mp3 from hf-internal-testing/dummy-audio-samples. Both policies preserve their source's active speaker-slot counts there. Full-recording fast32 still has 36 threshold crossings in a second slot where FP32 has none. These are source-agreement measurements, not human-annotated diarization error rates. FP16 is selected because INT8 introduced a substantial fast32 speaker-slot error without a useful throughput gain. Quality depends on the source model and the selected policy.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for coder543/nemotron-3-diarization-coreai

Finetuned
(7)
this model

Collection including coder543/nemotron-3-diarization-coreai