ReDimNet2 B3 (VoxCeleb2-dev, ptn) for MLX

MLX conversion of the official ReDimNet2 b3-vox2-ptn.pt checkpoint (pretrained, before large-margin fine-tuning). It produces 192-dimensional speaker embeddings from 16 kHz mono audio.

Usage

uv add git+https://github.com/avra-m3/speaker-embedding-mlx
import mlx.core as mx
from redimnet2_mlx import load_model

model = load_model("causal/redimnet2-b3-vox2-ptn-mlx")
emb = model(mx.array(wav_16k_float32)[None])  # (1, 192)

Parity

Embeddings match the PyTorch reference to 4.4e-06 relative error (worst case, float32, MLX CPU backend). See the parity log.

License and citation

MIT, following the upstream ReDimNet2 release (Copyright (c) 2026 Palabra.ai). See third-party notices.

@inproceedings{redimnet2_2026,
  title={ReDimNet2: Scaling Speaker Verification via Time-Pooled Dimension Reshaping},
  author={Yakovlev, Ivan and Okhotnikov, Anton},
  booktitle={Proceedings of Interspeech 2026},
  year={2026},
  eprint={2603.11841},
  archivePrefix={arXiv},
}
Downloads last month
18
MLX
Hardware compatibility
Log In to add your hardware

Quantized

Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Paper for causal/redimnet2-b3-vox2-ptn-mlx