ERes2NetV2 Chinese common for Sherpaw
Emscripten preload assets for local speaker embedding extraction with Sherpaw. The model consumes 16 kHz speech through sherpa-onnx and produces a 192-dimensional embedding. Speaker names and enrollment storage belong to the application.
Source and packaging
- Original model: ERes2NetV2 Chinese common, published by the 3D-Speaker / iic team under Apache-2.0.
- ONNX asset:
3dspeaker_speech_eres2netv2_sv_zh-cn_16k-common.onnx, distributed by sherpa-onnx. - ONNX SHA-256:
bf1a75b9930474cf3389ef415e6e5d38ca96fea4a3a00f7e301d080a58ee2239. - Preparation scripts: Sherpaw model preparation.
Weights are unchanged. Packaging renames the virtual file to /speaker-embedding.onnx and uses emscripten/emsdk:4.0.23. preload.data is byte-for-byte identical to the source ONNX file. manifest.json records source and artifact hashes.
Files and use
Download all three files from install/bin/wasm/:
preload.data: model bytes (71,441,526 bytes).preload.js: Emscripten ES module loader.preload.js.metadata: the virtual file name and byte range.
Use these with Sherpaw's separately built speaker-embedding WASM runtime and preload support. These are model assets, not a standalone runtime or a Transformers pipeline. Pin a Hugging Face commit when deploying, and serve the files where the loader can resolve them.
For consumers that load ONNX bytes directly, preload.data can also be used as the model byte buffer: this package contains exactly one ONNX file with no container header. Do not mix embeddings produced by different models, even if their dimensions match.
Reproduce
Run ./download.sh to fetch and verify the upstream ONNX asset, then ./pack.sh with Docker available to produce install/bin/wasm/. Model downloads are verified against the pinned checksum before replacing local files.
Limitations
Use single-speaker speech segments. Similarity is not a probability, and an acceptance threshold must be calibrated for the application. Unknown speakers, noisy or short recordings, and voice transformations can cause errors. This package does not establish accuracy or latency on a target device.
Model attribution and original documentation are linked above. See LICENSE for Apache-2.0 terms.