ERes2NetV2-192

Community Hugging Face Transformers packaging of the official 192-dimensional ERes2NetV2 speaker encoder from the 3D-Speaker project.

This repository is a community packaging. It is not an official Hugging Face, ModelScope, or 3D-Speaker repository.

Exact provenance

  • Upstream project: modelscope/3D-Speaker
  • Pinned 3D-Speaker commit: 065629c313eaf1a01c65c640c46d77e61e9607b4
  • Original ModelScope model: iic/speech_eres2netv2_sv_zh-cn_16k-common
  • Original ModelScope revision: v1.0.1
  • Original checkpoint: pretrained_eres2netv2.ckpt
  • Original checkpoint SHA-256: 0eb4057106b2573dd7b132cf0c36273ab29afd192c1610f80baa9c556dbb963c
  • Migrated model.safetensors SHA-256: 60f5928fce547c7353ba2609eb1aa9ceac7a8c1dd1e7e155e2be333c0aaf0c8c
  • Architecture: ERes2NetV2
  • Input: 16 kHz mono audio
  • Features: 80-dimensional Kaldi FBank, dither=0, per-utterance mean normalization
  • Embedding size: 192
  • License: Apache-2.0

The encoder tensors are loaded strictly from the original checkpoint and are not fine-tuned during migration.

Minimal source adaptation

The copied upstream core has only packaging-local import rewrites plus one numerically neutral autograd-safety change:

x = x.unsqueeze_(1)

→

x = x.unsqueeze(1)

This avoids an in-place view metadata mutation that breaks backward for cropped views from standard padded Transformers batches. No weights are changed.

Because ERes2NetV2 is not built into Transformers, use trust_remote_code=True.

Inference

import torch
from transformers import AutoFeatureExtractor, AutoModelForAudioXVector

repo_id = "artyomboyko/eres2netv2-192"
feature_extractor = AutoFeatureExtractor.from_pretrained(
    repo_id,
    trust_remote_code=True,
)
model = AutoModelForAudioXVector.from_pretrained(
    repo_id,
    trust_remote_code=True,
).eval()

inputs = feature_extractor(
    audio,
    sampling_rate=16000,
    return_tensors="pt",
)

with torch.inference_mode():
    outputs = model(**inputs)

embedding = outputs.embeddings  # [batch, 192], L2-normalized
raw_embedding = outputs.logits  # [batch, 192], raw encoder output

Variable-length batches use ordinary right padding:

inputs = feature_extractor(
    [audio_1, audio_2],
    sampling_rate=16000,
    padding=True,
    return_tensors="pt",
)
embeddings = model(**inputs).embeddings

Fine-tuning with standard Trainer

The pretrained repository contains only the speaker encoder. A classifier is created only when num_labels is supplied; the expected classifier.weight MISSING load message in that downstream case means the new task head is initialized from scratch.

from transformers import (
    AutoFeatureExtractor,
    AutoModelForAudioXVector,
    Trainer,
    TrainingArguments,
)

feature_extractor = AutoFeatureExtractor.from_pretrained(
    repo_id,
    trust_remote_code=True,
)
model = AutoModelForAudioXVector.from_pretrained(
    repo_id,
    num_labels=num_speakers,
    id2label=id2label,
    label2id=label2id,
    trust_remote_code=True,
)

training_args = TrainingArguments(
    output_dir="eres2netv2-finetuned",
    learning_rate=1e-5,
    per_device_train_batch_size=32,
    num_train_epochs=10,
)

trainer = Trainer(
    model=model,
    args=training_args,
    train_dataset=train_dataset,
    eval_dataset=eval_dataset,
    processing_class=feature_extractor,
)
trainer.train()

No custom Trainer, callback, or collator is required.

Training objective

When num_labels > 0, the wrapper uses:

  • CosineClassifier;
  • ArcMarginLoss;
  • scale 32.0;
  • margin 0.3;
  • easy_margin=False.

The classifier is not part of the original pretrained checkpoint.

Standard save / reload

model.save_pretrained("eres2netv2-finetuned")
feature_extractor.save_pretrained("eres2netv2-finetuned")

model = AutoModelForAudioXVector.from_pretrained(
    "eres2netv2-finetuned",
    trust_remote_code=True,
)

Scope

This repository provides speaker embeddings. It does not itself implement ASR, diarization, clustering, speaker enrollment policy, or online speaker memory.

References

Downloads last month
-
Safetensors
Model size
17.9M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support