Nemotron 3 Diarization, tuned decoder

NVIDIA's Nemotron 3 Diarization with a tuned decoder: 2.965% DER on the Omi clinical diarization benchmark, vs 4.789% for NVIDIA's default decoding. The weights are NVIDIA's, unmodified. What's new is how the frame probabilities become speaker segments: short pauses by the same speaker are bridged instead of split.

Results

Omi clinical diarization benchmark: 15 simulated doctor-patient consultations from PriMock57, 2.4 hours, automatic speaker count. DER at ±250 ms uses the benchmark's own scorer; strict is zero tolerance.

System DER ±250 ms ↓ Strict DER ↓ Correct speaker count
Pyannote Precision-3 API 2.891% not published 14/15
This model 2.965% 12.942% 14/15
Nemotron 3 + Omi runtime (proprietary) 3.174% 13.203% 14/15
Nemotron 3, NVIDIA default decoding 4.803% 12.720% 14/15
Pyannote Community-1 6.620% not published 9/15
Sortformer v2.1 7.974% not published 3/15

Competitor rows are Omi's published numbers. This model's row was scored from FP32 CPU probabilities; decoding those with NVIDIA's default 0.5 threshold gives 4.789%, matching NVIDIA's published 4.803% output at 99.8% frame agreement.

Caveats

  • The decoder settings were chosen on these same 15 recordings, as Omi's were. Expect a smaller gain on new audio.
  • The recordings are simulated two-speaker consultations. Nothing here tests 3+ speakers, noisy rooms or real clinics.
  • Bridging pauses trades a little strict accuracy for fewer broken turns; at ±250 ms it wins, at zero tolerance this model is 0.2 pp worse than NVIDIA's default.

How the decoder works

  1. Speaker active when its probability is above 0.5.
  2. A same-speaker pause under 0.4 s is filled.
  3. A pause under 0.5 s is filled if no other speaker talks during it.

Settings live in decoder_config.json. Decoding takes about 1 ms per recording.

Usage

pip install torch soundfile huggingface_hub git+https://github.com/huggingface/transformers
import sys
from huggingface_hub import snapshot_download

path = snapshot_download("aduomas/nemotron-3-diarization")
sys.path.insert(0, path)
from diarize import Diarizer

diarizer = Diarizer(path)
for segment in diarizer.diarize("visit.wav"):  # 16 kHz mono
    print(segment)  # {"start": 1.42, "end": 1.82, "speaker": "spk_0"}

On a CUDA GPU the model runs in BF16, and the per-chunk step is compiled into CUDA graphs. Omi reported that approach roughly halving inference time on an L4 (0.674 s to 0.323 s per ~10-minute file); it is not yet measured for this release. Diarizer.probabilities() returns the raw (frames, 8) probabilities at 10 ms.

Licence and attribution

Model weights, config.json and processor_config.json are from nvidia/Nemotron-3-Diarization (revision a435e98), redistributed under the OpenMDW License 1.1. See NOTICE. Refer to NVIDIA's model card for training data, intended use, bias and safety information.

Downloads last month
22
Safetensors
Model size
99.2M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Aduomas/nemotron-3-diarization

Finetuned
(9)
this model