Update README.md
Browse files
README.md
CHANGED
@@ -31,7 +31,7 @@ model-index:
|
|
31 |
|
32 |
# Wav2Vec2-Large-XLSR-53-ml
|
33 |
|
34 |
-
Fine-tuned [facebook/wav2vec2-large-xlsr-53](https://huggingface.co/facebook/wav2vec2-large-xlsr-53) on ml (Malayalam) using the [Indic TTS Malayalam Speech Corpus (via Kaggle)](https://www.kaggle.com/kavyamanohar/indic-tts-malayalam-speech-corpus), [Openslr Malayalam Speech Corpus](http://openslr.org/63/), [SMC Malayalam Speech Corpus](https://blog.smc.org.in/malayalam-speech-corpus/) and [IIIT-H Indic Speech Databases](http://speech.iiit.ac.in/index.php/research-svl/69.html). The notebooks used to train model
|
35 |
|
36 |
## Usage
|
37 |
|
@@ -134,8 +134,8 @@ resamplers = {
|
|
134 |
48000: torchaudio.transforms.Resample(48_000, 16_000),
|
135 |
}
|
136 |
|
137 |
-
chars_to_ignore_regex = '[
|
138 |
-
unicode_ignore_regex = r'[
|
139 |
|
140 |
# Preprocessing the datasets.
|
141 |
# We need to read the audio files as arrays
|
|
|
31 |
|
32 |
# Wav2Vec2-Large-XLSR-53-ml
|
33 |
|
34 |
+
Fine-tuned [facebook/wav2vec2-large-xlsr-53](https://huggingface.co/facebook/wav2vec2-large-xlsr-53) on ml (Malayalam) using the [Indic TTS Malayalam Speech Corpus (via Kaggle)](https://www.kaggle.com/kavyamanohar/indic-tts-malayalam-speech-corpus), [Openslr Malayalam Speech Corpus](http://openslr.org/63/), [SMC Malayalam Speech Corpus](https://blog.smc.org.in/malayalam-speech-corpus/) and [IIIT-H Indic Speech Databases](http://speech.iiit.ac.in/index.php/research-svl/69.html). The notebooks used to train model are available [here](https://github.com/gauthamsuresh09/wav2vec2-large-xlsr-53-malayalam/). When using this model, make sure that your speech input is sampled at 16kHz.
|
35 |
|
36 |
## Usage
|
37 |
|
|
|
134 |
48000: torchaudio.transforms.Resample(48_000, 16_000),
|
135 |
}
|
136 |
|
137 |
+
chars_to_ignore_regex = '[\\\\,\\\\?\\\\.\\\\!\\\\-\\\\;\\\\:\\\\"\\\\“\\\\%\\\\‘\\\\”\\\\�Utrnle\\\\_]'
|
138 |
+
unicode_ignore_regex = r'[\\\\u200e]'
|
139 |
|
140 |
# Preprocessing the datasets.
|
141 |
# We need to read the audio files as arrays
|