DiariZen v2 for CoreML
CoreML conversions of the models behind DiariZen's
wavlm-large-s80-md-v2 speaker-diarization pipeline, packaged for Apple silicon. Used by the
transcriber desktop app through FluidAudio's
offline diarization pipeline. Every conversion is an exact algebraic rewrite of the source
checkpoint (reshapes, framing, normalization identities, activation identities); nothing was
retrained, fine-tuned or further pruned.
| file | source checkpoint | interface | placement |
|---|---|---|---|
DiariZenSegmentation.mlmodelc |
BUT-FIT/diarizen-wavlm-large-s80-md-v2 (pruned WavLM-Large + Conformer, 4 local speakers, 16-class powerset) | waveform (1, 256000) float32, 16 kHz mono โ logprobs (1, 799, 16) |
Neural Engine (1253 of 1255 ops) |
DiariZenFBank.mlmodelc |
Kaldi fbank + per-window mean normalization as used by pyannote's WeSpeaker wrapper | audio (B, 1, 256000), B in 1โฆ32 โ fbank_features (B, 1, 80, 1598) |
CPU |
DiariZenEmbedding.mlmodelc |
pyannote/wespeaker-voxceleb-resnet34-LM | fbank_features (1, 1, 80, 1598) + weights (4, 799) mask lanes โ embedding (4, 256) |
Neural Engine |
PldaRho.mlmodelc, plda-parameters.json, xvector-transform.json |
DiariZen's PLDA (plda.npz, xvec_transform.npz), byte-identical to pyannote community-1's |
256 โ 128 | CPU |
The embedding model runs the ResNet trunk once per 16 s window and pools up to four speaker masks from it, which is what DiariZen's per-slot embedding computes, at a quarter of the calls.
Verified against the PyTorch pipeline on ICSI meetings: segmentation log-probabilities within fp16 rounding (argmax agreement โฅ 99.75 %), embeddings cosine โฅ 0.99997, and identical per-embedding cluster labels through PLDA + AHC + VBx; meeting-level DER equal to the reference.
Requirements: macOS 14 or later, Apple silicon. First load compiles the segmentation model for the Neural Engine (10โ25 s, cached afterwards).
Licenses
- The segmentation model is a derivative of DiariZen's weights, released under CC BY-NC 4.0 by BUT-FIT. This repository is therefore CC BY-NC 4.0: attribution required, non-commercial use only.
- The fbank and embedding models derive from
pyannote/wespeaker-voxceleb-resnet34-LM(CC BY 4.0), itself a port of WeSpeaker's VoxCeleb ResNet34-LM. - PLDA parameters are DiariZen's, distributed with the checkpoint above.
Please cite DiariZen if you use these models:
Jiangyu Han et al., "Leveraging Self-Supervised Learning for Speaker Diarization" and "Fine-tuning Self-Supervised Models for Speaker Diarization with Structured Pruning" (BUT-FIT, 2025).
Model tree for sl-data-systems/diarizen-v2-coreml
Base model
BUT-FIT/diarizen-wavlm-large-s80-md-v2