SenseVoiceSmall:显式 Transformers 实现

本仓库将原始 SenseVoiceSmall 转换为标准 HF 配置与 FP32 safetensors。 不依赖 FunASR 或 ModelScope,不使用 AutoModel、模型注册表或远端代码执行。 原始模型与代码的来源、授权分别见 MODEL_LICENSE.txt 和 UPSTREAM-CODE-MIT.txt。 转换不训练、不量化、不改变参数:917 个参数张量逐项按位一致。

安装与下载

pip install torch==2.8.0 torchaudio==2.8.0 transformers==4.57.6 sentencepiece numpy safetensors soundfile
hf download MigoXV/SenseVoiceSmall --local-dir ./SenseVoiceSmall

直接使用具体模型类

仓库附带可独立导入的 sensevoice_hf 源码包。将下载目录加入 Python 路径,显式导入:

import sys
import soundfile as sf
import torch
sys.path.insert(0, "./SenseVoiceSmall")
from sensevoice_hf import SenseVoiceSmall, SenseVoiceProcessor

processor = SenseVoiceProcessor.from_pretrained("./SenseVoiceSmall", local_files_only=True)
model = SenseVoiceSmall.from_pretrained("./SenseVoiceSmall", local_files_only=True).eval()
audio, sr = sf.read("audio.wav", dtype="float32")
inputs = processor(audio, sampling_rate=sr, language="auto", use_itn=True)
with torch.inference_mode():
    output = model(**inputs)
results = processor.decode(output.logits, output.output_lengths)
print(results[0]["text"])

音频使用浮点幅度,单声道为一维,多声道为 channels-first; SoundFile 返回的多声道数据应先转置。批量输入使用一维音频数组的列表。 语言支持 auto/zh/en/yue/ja/ko/nospeech。模型是完整音频段的非自回归 CTC 模型。

亦可从已安装的 sensevoice-asr 项目导入具体类,并显式调用 SenseVoiceSmall.from_pretrained("MigoXV/SenseVoiceSmall", revision="指定提交")。 无需 trust_remote_code,不注册 Auto 类。

配置、词表和 CMVN 与权重一起保存;完整转换校验信息见 conversion_manifest.json。 CPU/GPU 共64个场景、648个阶段对齐全部通过,最大绝对差为0;权重917个张量按位一致。 源音频和测试转写不随模型上传,公开统计见 alignment_summary.json。

Downloads last month
14
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for MigoXV/SenseVoiceSmall

Finetuned
(22)
this model
Finetunes
1 model