AndrewMcDowell
/

wav2vec2-xls-r-300m-japanese

Automatic Speech Recognition

Generated from Trainer

hf-asr-leaderboard

mozilla-foundation/common_voice_8_0

robust-speech-event

Inference Endpoints

Model card Files Files and versions Community

AndrewMcDowell commited on Feb 4, 2022

Commit

504d95e

•

1 Parent(s): bf0f232

Update README.md

Add eval results.

Files changed (1) hide show

README.md +27 -2

README.md CHANGED Viewed

@@ -9,8 +9,23 @@ tags:
 datasets:
 - common_voice
 model-index:
-- name: ''
-  results: []
 ---
 <!-- This model card has been generated automatically according to the information the Trainer had access to. You
@@ -19,6 +34,9 @@ should probably proofread and complete it, then remove this comment. -->
 #
 This model is a fine-tuned version of [facebook/wav2vec2-xls-r-300m](https://huggingface.co/facebook/wav2vec2-xls-r-300m) on the MOZILLA-FOUNDATION/COMMON_VOICE_8_0 - JA dataset.
 It achieves the following results on the evaluation set:
 - Loss: 0.5212
 - Wer: 1.3068
@@ -72,3 +90,10 @@ The following hyperparameters were used during training:
 - Pytorch 1.10.2+cu102
 - Datasets 1.18.2.dev0
 - Tokenizers 0.11.0

 datasets:
 - common_voice
 model-index:
+- name: 'XLS-R-300-m'
+  results:
+  - task:
+      name: Automatic Speech Recognition
+      type: automatic-speech-recognition
+    dataset:
+      name: Common Voice 8
+      type: mozilla-foundation/common_voice_8_0
+      args: ja
+    metrics:
+       - name: Test WER
+         type: wer
+         value: 95.82
+       - name: Test CER
+         type: cer
+         value: 23.64
 ---
 <!-- This model card has been generated automatically according to the information the Trainer had access to. You
 #
 This model is a fine-tuned version of [facebook/wav2vec2-xls-r-300m](https://huggingface.co/facebook/wav2vec2-xls-r-300m) on the MOZILLA-FOUNDATION/COMMON_VOICE_8_0 - JA dataset.
+Kanji are converted into Hiragana using the [pykakasi](https://pykakasi.readthedocs.io/en/latest/index.html) library during training and evaluation. The model can output both Hiragana and Katakana characters.
 It achieves the following results on the evaluation set:
 - Loss: 0.5212
 - Wer: 1.3068
 - Pytorch 1.10.2+cu102
 - Datasets 1.18.2.dev0
 - Tokenizers 0.11.0
+#### Evaluation Commands
+1. To evaluate on `mozilla-foundation/common_voice_8_0` with split `test`
+```bash
+python ./eval.py --model_id AndrewMcDowell/wav2vec2-xls-r-300m-japanese --dataset mozilla-foundation/common_voice_8_0 --config ja --split test --log_outputs
+```