FastConformer Hybrid Large PC (French), CTC head, sherpa-onnx int8

The French speech recognition model that the «Слышно» Android app downloads for on-device transcription. This repository is the fallback source for the app's own mirror; the app downloads every file separately and verifies its size and SHA-256.

Attribution

  • Original model: nvidia/stt_fr_fastconformer_hybrid_large_pc, © NVIDIA Corporation, licensed under CC BY 4.0.
  • Changes made:
    1. The encoder with the CTC head was exported to ONNX by OpenVoiceOS: OpenVoiceOS/stt_fr_fastconformer_hybrid_large_pc_onnx (model.onnx, vocab.txt, CC BY 4.0).
    2. For the Slyshno app, sherpa-onnx metadata was added to that graph (vocab_size = 1025, normalize_type = per_feature, subsampling_factor = 8, model_type = EncDecHybridRNNTCTCBPEModel, language = fr), and the graph was dynamically quantized to int8 with ONNX Runtime — MatMul nodes only, uint8 weights; convolutions stay float. tokens.txt is the upstream vocab.txt, unchanged.

Files

file bytes sha256
model.int8.onnx 173888277 11dd49f5d63cf948f982a1a95a68d018e3de10a83f3bba642e16db73989d3e43
tokens.txt 10943 1b0a63466e3847115896e795fd64160e689c68efb63c0b648f7a85d33a9a0c1e

Usage

sherpa-onnx OfflineRecognizer with a NeMo CTC config (nemo_ctc.model = model.int8.onnx, tokens = tokens.txt), 16 kHz mono input, 80-dim features. The output carries punctuation and capitalisation.

Accuracy

FLEURS fr_fr dev, sherpa-onnx 1.13.8, 2 threads, case and punctuation removed before scoring: first 12 clips WER 9.1 %, CER 3.9 %; first 100 clips WER 9.0 %, CER 4.2 %.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for krut42/voice-fastconformer-fr-ctc-int8

Quantized
(3)
this model