YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

SmartTTS β€” Indian Language Identification Model

Model Description

A CNN-LSTM hybrid deep learning model trained to identify spoken Indian languages from audio files. Built as part of the SmartTTS project β€” an end-to-end multilingual text-to-speech pipeline.

Training Data

  • Dataset: Indian Languages Audio Dataset
  • Total samples: 10,013 MP3 files
  • Duration: 5 seconds per file
  • Languages: Bengali, English, Gujarati, Hindi, Kannada, Malayalam, Marathi, Punjabi, Tamil, Telugu, Urdu

Model Architecture

  • 6 Γ— Conv1D layers (pattern detection)
  • 3 Γ— MaxPooling layers (dimensionality reduction)
  • 2 Γ— LSTM layers (temporal sequence memory)
  • 2 Γ— Dense layers (classification)
  • Output: Softmax over 11 language classes

Features Extracted (194 total)

  • MFCC (40) + Delta MFCC (40)
  • Chroma (12) + Mel Spectrogram (40)
  • Spectral Contrast (7) + ZCR + Rolloff + RMS

Performance

Metric Value
Overall Test Accuracy 87.55%
Bengali 99%
English 99%
Gujarati 99%
Malayalam 99%
Marathi 99%
Hindi 97%
Kannada 97%
Tamil 97%
Punjabi 98%
Telugu 93%

How This Model Is Used

This model acts as a quality verification layer in the SmartTTS pipeline:

  1. User selects a language and generates speech via Edge TTS
  2. This CNN-LSTM model analyses the generated audio
  3. Returns confidence score confirming the audio matches the selected language
  4. Result is displayed live on the frontend as a verification badge

Live Demo

Try it at: SmartTTS Space

Citation

@misc{smarttts2025,
  author = {Kartavya11},
  title = {SmartTTS Indian Language Identification},
  year = {2025},
  publisher = {HuggingFace}
}
Downloads last month
13
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Space using Kartavya11/language-detection 1