ReDimNet2: Scaling Speaker Verification via Time-Pooled Dimension Reshaping
Paper • 2603.11841 • Published • 3
How to use causal/redimnet2-b3-vox2-ptn-mlx with MLX:
# Download the model from the Hub pip install huggingface_hub[hf_xet] hf download causal/redimnet2-b3-vox2-ptn-mlx --local-dir redimnet2-b3-vox2-ptn-mlx
MLX conversion of the official ReDimNet2 b3-vox2-ptn.pt checkpoint
(pretrained, before large-margin fine-tuning). It produces 192-dimensional speaker
embeddings from 16 kHz mono audio.
uv add git+https://github.com/avra-m3/speaker-embedding-mlx
import mlx.core as mx
from redimnet2_mlx import load_model
model = load_model("causal/redimnet2-b3-vox2-ptn-mlx")
emb = model(mx.array(wav_16k_float32)[None]) # (1, 192)
Embeddings match the PyTorch reference to 4.4e-06 relative error (worst case, float32, MLX CPU backend). See the parity log.
MIT, following the upstream ReDimNet2 release (Copyright (c) 2026 Palabra.ai). See third-party notices.
@inproceedings{redimnet2_2026,
title={ReDimNet2: Scaling Speaker Verification via Time-Pooled Dimension Reshaping},
author={Yakovlev, Ivan and Okhotnikov, Anton},
booktitle={Proceedings of Interspeech 2026},
year={2026},
eprint={2603.11841},
archivePrefix={arXiv},
}
Quantized