Instructions to use krut42/voice-fastconformer-fr-ctc-int8 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- NeMo
How to use krut42/voice-fastconformer-fr-ctc-int8 with NeMo:
import nemo.collections.asr as nemo_asr asr_model = nemo_asr.models.ASRModel.from_pretrained("krut42/voice-fastconformer-fr-ctc-int8") transcriptions = asr_model.transcribe(["file.wav"]) - Notebooks
- Google Colab
- Kaggle
FastConformer Hybrid Large PC (French), CTC head, sherpa-onnx int8
The French speech recognition model that the «Слышно» Android app downloads for on-device transcription. This repository is the fallback source for the app's own mirror; the app downloads every file separately and verifies its size and SHA-256.
Attribution
- Original model: nvidia/stt_fr_fastconformer_hybrid_large_pc, © NVIDIA Corporation, licensed under CC BY 4.0.
- Changes made:
- The encoder with the CTC head was exported to ONNX by OpenVoiceOS:
OpenVoiceOS/stt_fr_fastconformer_hybrid_large_pc_onnx
(
model.onnx,vocab.txt, CC BY 4.0). - For the Slyshno app, sherpa-onnx metadata was added to that graph
(
vocab_size = 1025,normalize_type = per_feature,subsampling_factor = 8,model_type = EncDecHybridRNNTCTCBPEModel,language = fr), and the graph was dynamically quantized to int8 with ONNX Runtime — MatMul nodes only, uint8 weights; convolutions stay float.tokens.txtis the upstreamvocab.txt, unchanged.
- The encoder with the CTC head was exported to ONNX by OpenVoiceOS:
OpenVoiceOS/stt_fr_fastconformer_hybrid_large_pc_onnx
(
Files
| file | bytes | sha256 |
|---|---|---|
model.int8.onnx |
173888277 | 11dd49f5d63cf948f982a1a95a68d018e3de10a83f3bba642e16db73989d3e43 |
tokens.txt |
10943 | 1b0a63466e3847115896e795fd64160e689c68efb63c0b648f7a85d33a9a0c1e |
Usage
sherpa-onnx OfflineRecognizer with a NeMo CTC config (nemo_ctc.model = model.int8.onnx,
tokens = tokens.txt), 16 kHz mono input, 80-dim features. The output carries
punctuation and capitalisation.
Accuracy
FLEURS fr_fr dev, sherpa-onnx 1.13.8, 2 threads, case and punctuation removed before
scoring: first 12 clips WER 9.1 %, CER 3.9 %; first 100 clips WER 9.0 %,
CER 4.2 %.
Model tree for krut42/voice-fastconformer-fr-ctc-int8
Base model
nvidia/stt_fr_fastconformer_hybrid_large_pc