Automatic Speech Recognition
Transformers
Safetensors
sensevoice_small
sensevoice
speech-recognition
Instructions to use MigoXV/SenseVoiceSmall with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use MigoXV/SenseVoiceSmall with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("automatic-speech-recognition", model="MigoXV/SenseVoiceSmall")# pip install -U transformers accelerate # Load model directly from transformers import SenseVoiceSmall model = SenseVoiceSmall.from_pretrained("MigoXV/SenseVoiceSmall", device_map="auto") - Notebooks
- Google Colab
- Kaggle
SenseVoiceSmall:显式 Transformers 实现
本仓库将原始 SenseVoiceSmall 转换为标准 HF 配置与 FP32 safetensors。 不依赖 FunASR 或 ModelScope,不使用 AutoModel、模型注册表或远端代码执行。 原始模型与代码的来源、授权分别见 MODEL_LICENSE.txt 和 UPSTREAM-CODE-MIT.txt。 转换不训练、不量化、不改变参数:917 个参数张量逐项按位一致。
安装与下载
pip install torch==2.8.0 torchaudio==2.8.0 transformers==4.57.6 sentencepiece numpy safetensors soundfile
hf download MigoXV/SenseVoiceSmall --local-dir ./SenseVoiceSmall
直接使用具体模型类
仓库附带可独立导入的 sensevoice_hf 源码包。将下载目录加入 Python 路径,显式导入:
import sys
import soundfile as sf
import torch
sys.path.insert(0, "./SenseVoiceSmall")
from sensevoice_hf import SenseVoiceSmall, SenseVoiceProcessor
processor = SenseVoiceProcessor.from_pretrained("./SenseVoiceSmall", local_files_only=True)
model = SenseVoiceSmall.from_pretrained("./SenseVoiceSmall", local_files_only=True).eval()
audio, sr = sf.read("audio.wav", dtype="float32")
inputs = processor(audio, sampling_rate=sr, language="auto", use_itn=True)
with torch.inference_mode():
output = model(**inputs)
results = processor.decode(output.logits, output.output_lengths)
print(results[0]["text"])
音频使用浮点幅度,单声道为一维,多声道为 channels-first; SoundFile 返回的多声道数据应先转置。批量输入使用一维音频数组的列表。 语言支持 auto/zh/en/yue/ja/ko/nospeech。模型是完整音频段的非自回归 CTC 模型。
亦可从已安装的 sensevoice-asr 项目导入具体类,并显式调用
SenseVoiceSmall.from_pretrained("MigoXV/SenseVoiceSmall", revision="指定提交")。
无需 trust_remote_code,不注册 Auto 类。
配置、词表和 CMVN 与权重一起保存;完整转换校验信息见 conversion_manifest.json。 CPU/GPU 共64个场景、648个阶段对齐全部通过,最大绝对差为0;权重917个张量按位一致。 源音频和测试转写不随模型上传,公开统计见 alignment_summary.json。
- Downloads last month
- 10