indic-asr-multi

Multilingual speech recognition for Indian languages. This is a Conformer RNN-T model fine-tuned on Indian-language speech. During fine-tuning it also learned to recognise which language is being spoken, so no language code is needed. It outputs text in the native script of the detected language.

Languages: Hindi, Marathi, Telugu, Tamil, Kannada, Malayalam, Gujarati, Punjabi, Odia

Usage

pip install "nemo_toolkit[asr]"
import nemo.collections.asr as nemo_asr

model = nemo_asr.models.ASRModel.from_pretrained("Archit-01/indic-asr-multi")
out = model.transcribe(["audio.wav"])

if isinstance(out, tuple):  # older NeMo versions
    out = out[0]
print(out[0].text if hasattr(out[0], "text") else out[0])

Input: mono WAV audio. The model runs at 16000 Hz, and files at other sample rates are resampled automatically.

Limitations

  • Very short clips can occasionally be decoded in the wrong language's script.
  • Accuracy drops on rare words, names, and heavily code-mixed speech.
Downloads last month
2
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support