Instructions to use Aduomas/nemotron-3-diarization with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Aduomas/nemotron-3-diarization with Transformers:
# Load model directly from transformers import AutoProcessor, AutoModelForAudioFrameClassification processor = AutoProcessor.from_pretrained("Aduomas/nemotron-3-diarization") model = AutoModelForAudioFrameClassification.from_pretrained("Aduomas/nemotron-3-diarization", device_map="auto") - Notebooks
- Google Colab
- Kaggle
Nemotron 3 Diarization, tuned decoder
NVIDIA's Nemotron 3 Diarization with a tuned decoder: 2.965% DER on the Omi clinical diarization benchmark, vs 4.789% for NVIDIA's default decoding. The weights are NVIDIA's, unmodified. What's new is how the frame probabilities become speaker segments: short pauses by the same speaker are bridged instead of split.
Results
Omi clinical diarization benchmark: 15 simulated doctor-patient consultations from PriMock57, 2.4 hours, automatic speaker count. DER at ±250 ms uses the benchmark's own scorer; strict is zero tolerance.
| System | DER ±250 ms ↓ | Strict DER ↓ | Correct speaker count |
|---|---|---|---|
| Pyannote Precision-3 API | 2.891% | not published | 14/15 |
| This model | 2.965% | 12.942% | 14/15 |
| Nemotron 3 + Omi runtime (proprietary) | 3.174% | 13.203% | 14/15 |
| Nemotron 3, NVIDIA default decoding | 4.803% | 12.720% | 14/15 |
| Pyannote Community-1 | 6.620% | not published | 9/15 |
| Sortformer v2.1 | 7.974% | not published | 3/15 |
Competitor rows are Omi's published numbers. This model's row was scored from FP32 CPU probabilities; decoding those with NVIDIA's default 0.5 threshold gives 4.789%, matching NVIDIA's published 4.803% output at 99.8% frame agreement.
Caveats
- The decoder settings were chosen on these same 15 recordings, as Omi's were. Expect a smaller gain on new audio.
- The recordings are simulated two-speaker consultations. Nothing here tests 3+ speakers, noisy rooms or real clinics.
- Bridging pauses trades a little strict accuracy for fewer broken turns; at ±250 ms it wins, at zero tolerance this model is 0.2 pp worse than NVIDIA's default.
How the decoder works
- Speaker active when its probability is above 0.5.
- A same-speaker pause under 0.4 s is filled.
- A pause under 0.5 s is filled if no other speaker talks during it.
Settings live in decoder_config.json. Decoding takes about 1 ms per recording.
Usage
pip install torch soundfile huggingface_hub git+https://github.com/huggingface/transformers
import sys
from huggingface_hub import snapshot_download
path = snapshot_download("aduomas/nemotron-3-diarization")
sys.path.insert(0, path)
from diarize import Diarizer
diarizer = Diarizer(path)
for segment in diarizer.diarize("visit.wav"): # 16 kHz mono
print(segment) # {"start": 1.42, "end": 1.82, "speaker": "spk_0"}
On a CUDA GPU the model runs in BF16, and the per-chunk step is compiled into CUDA graphs. Omi reported that
approach roughly halving inference time on an L4 (0.674 s to 0.323 s per ~10-minute file); it is not yet
measured for this release. Diarizer.probabilities() returns the raw (frames, 8) probabilities at 10 ms.
Licence and attribution
Model weights, config.json and processor_config.json are from
nvidia/Nemotron-3-Diarization (revision a435e98),
redistributed under the OpenMDW License 1.1. See NOTICE. Refer to NVIDIA's model card for training
data, intended use, bias and safety information.
- Downloads last month
- 22
Model tree for Aduomas/nemotron-3-diarization
Base model
nvidia/Nemotron-3-Diarization