Speaker Diarization for Core ML
Who spoke when, on Apple Silicon. These are the pyannote speaker-diarization community-1 models (powerset speaker segmentation, WeSpeaker ResNet speaker embeddings, PLDA scoring for VBx clustering) converted to Core ML and laid out for the Neural Engine. They work on any language because they model the voice, not the words.
This repository is a pinned copy of
FluidInference/speaker-diarization-coreml
at revision 1ed7a662fdc7109e36d822db793ee6eebdaf8594. The files are byte-identical to that revision; only
this card differs.
Requirements
- macOS 14 or later, or iOS 17 or later
- Apple Silicon recommended
Download
hf download 1of2/speaker-diarization-coreml --local-dir speaker-diarization-coreml
Models
All models take 16 kHz mono Float32 audio in 10 s windows (160,000 samples). Segmentation produces 589 frames per window (about 17 ms each).
Streaming pipeline
| Model | Inputs | Outputs |
|---|---|---|
pyannote_segmentation.mlmodelc |
audio [1, 1, 160000] |
segments [1, 589, 7]: powerset log-probabilities for up to 3 local speakers |
wespeaker_v2.mlmodelc |
waveform [3, 160000], mask [3, 589] |
embedding [3, 256]: one speaker vector per local speaker |
Slide the window, decode the powerset into per-speaker activity, embed each active speaker with its mask, and link speakers across windows by cosine similarity of the 256-dim embeddings.
Offline pipeline (VBx)
| Model / file | Inputs | Outputs |
|---|---|---|
Segmentation.mlmodelc |
audio [32, 1, 160000] (batch of 32 windows) |
log_probs |
FBank.mlmodelc |
audio [1, 1, 160000] |
fbank_features (80 mel bins) |
Embedding.mlmodelc |
fbank_features [1, 1, 80, 998], weights [1, 589] |
embedding (256-dim) |
PldaRho.mlmodelc |
embeddings [32, 256] |
rho (PLDA-space features) |
plda-parameters.json, xvector-transform.json |
PLDA and x-vector transform parameters for clustering |
Batch segmentation over the whole recording, embed each speaker-weighted window, project into PLDA space, then cluster with Variational Bayes HMM (VBx). This gives the best accuracy on complete recordings.
mlpackages/ holds the source packages; the remaining .mlmodelc files are earlier exports kept for
compatibility.
Accuracy and speed
The conversion tracks the PyTorch reference closely: end-to-end DER and JER stay within about 1% of the original pipeline. Core ML runs roughly 10x faster than PyTorch on CPU and 20x on GPU.
Small precision differences in segmentation scores are absorbed by clustering:
See the original model for DER benchmarks.
Limitations
- At most 3 speakers active within a single 10 s window; the number of speakers per recording is unbounded.
- Heavily overlapped speech and very short turns (under about half a second) are the main error sources.
- Speaker labels are anonymous; naming them needs enrolment embeddings supplied by the caller.
License and attribution
Licensed under CC BY 4.0.
- Original models: pyannote speaker-diarization-community-1, CC BY 4.0.
- Core ML conversion: FluidInference.
@inproceedings{Plaquet23,
author={Alexis Plaquet and Hervé Bredin},
title={{Powerset multi-class cross entropy loss for neural speaker diarization}},
year=2023,
booktitle={Proc. INTERSPEECH 2023},
}
@inproceedings{Wang2023,
title={Wespeaker: A research and production oriented speaker embedding learning toolkit},
author={Wang, Hongji and Liang, Chengdong and Wang, Shuai and Chen, Zhengyang and Zhang, Binbin and Xiang, Xu and Deng, Yanlei and Qian, Yanmin},
booktitle={ICASSP 2023, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
pages={1--5},
year={2023},
}
@article{Landini2022,
author={Landini, Federico and Profant, J{\'a}n and Diez, Mireia and Burget, Luk{\'a}{\v{s}}},
title={{Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: theory, implementation and analysis on standard tasks}},
year={2022},
journal={Computer Speech \& Language},
}
- Downloads last month
- -
Model tree for 1of2/speaker-diarization-coreml
Base model
pyannote/speaker-diarization-community-1

