Speaker Diarization for Core ML

Who spoke when, on Apple Silicon. These are the pyannote speaker-diarization community-1 models (powerset speaker segmentation, WeSpeaker ResNet speaker embeddings, PLDA scoring for VBx clustering) converted to Core ML and laid out for the Neural Engine. They work on any language because they model the voice, not the words.

This repository is a pinned copy of FluidInference/speaker-diarization-coreml at revision 1ed7a662fdc7109e36d822db793ee6eebdaf8594. The files are byte-identical to that revision; only this card differs.

Requirements

  • macOS 14 or later, or iOS 17 or later
  • Apple Silicon recommended

Download

hf download 1of2/speaker-diarization-coreml --local-dir speaker-diarization-coreml

Models

All models take 16 kHz mono Float32 audio in 10 s windows (160,000 samples). Segmentation produces 589 frames per window (about 17 ms each).

Streaming pipeline

Model Inputs Outputs
pyannote_segmentation.mlmodelc audio [1, 1, 160000] segments [1, 589, 7]: powerset log-probabilities for up to 3 local speakers
wespeaker_v2.mlmodelc waveform [3, 160000], mask [3, 589] embedding [3, 256]: one speaker vector per local speaker

Slide the window, decode the powerset into per-speaker activity, embed each active speaker with its mask, and link speakers across windows by cosine similarity of the 256-dim embeddings.

Offline pipeline (VBx)

Model / file Inputs Outputs
Segmentation.mlmodelc audio [32, 1, 160000] (batch of 32 windows) log_probs
FBank.mlmodelc audio [1, 1, 160000] fbank_features (80 mel bins)
Embedding.mlmodelc fbank_features [1, 1, 80, 998], weights [1, 589] embedding (256-dim)
PldaRho.mlmodelc embeddings [32, 256] rho (PLDA-space features)
plda-parameters.json, xvector-transform.json PLDA and x-vector transform parameters for clustering

Batch segmentation over the whole recording, embed each speaker-weighted window, project into PLDA space, then cluster with Variational Bayes HMM (VBx). This gives the best accuracy on complete recordings.

mlpackages/ holds the source packages; the remaining .mlmodelc files are earlier exports kept for compatibility.

Accuracy and speed

The conversion tracks the PyTorch reference closely: end-to-end DER and JER stay within about 1% of the original pipeline. Core ML runs roughly 10x faster than PyTorch on CPU and 20x on GPU.

Pipeline timing

Pipeline overview

Small precision differences in segmentation scores are absorbed by clustering:

Metric drift

See the original model for DER benchmarks.

Limitations

  • At most 3 speakers active within a single 10 s window; the number of speakers per recording is unbounded.
  • Heavily overlapped speech and very short turns (under about half a second) are the main error sources.
  • Speaker labels are anonymous; naming them needs enrolment embeddings supplied by the caller.

License and attribution

Licensed under CC BY 4.0.

@inproceedings{Plaquet23,
  author={Alexis Plaquet and Hervé Bredin},
  title={{Powerset multi-class cross entropy loss for neural speaker diarization}},
  year=2023,
  booktitle={Proc. INTERSPEECH 2023},
}

@inproceedings{Wang2023,
  title={Wespeaker: A research and production oriented speaker embedding learning toolkit},
  author={Wang, Hongji and Liang, Chengdong and Wang, Shuai and Chen, Zhengyang and Zhang, Binbin and Xiang, Xu and Deng, Yanlei and Qian, Yanmin},
  booktitle={ICASSP 2023, IEEE International Conference on Acoustics, Speech and Signal Processing (ICASSP)},
  pages={1--5},
  year={2023},
}

@article{Landini2022,
  author={Landini, Federico and Profant, J{\'a}n and Diez, Mireia and Burget, Luk{\'a}{\v{s}}},
  title={{Bayesian HMM clustering of x-vector sequences (VBx) in speaker diarization: theory, implementation and analysis on standard tasks}},
  year={2022},
  journal={Computer Speech \& Language},
}
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for 1of2/speaker-diarization-coreml

Quantized
(6)
this model