AndrewMcDowell
/

wav2vec2-xls-r-300m-japanese

Automatic Speech Recognition

Generated from Trainer

hf-asr-leaderboard

mozilla-foundation/common_voice_8_0

robust-speech-event

Inference Endpoints

Model card Files Files and versions Community

AndrewMcDowell commited on Feb 3, 2022

Commit

b4be586

·

1 Parent(s): 6767044

Update README.md

Add eval metrics.

Files changed (1) hide show

README.md +27 -2

README.md CHANGED Viewed

@@ -6,11 +6,27 @@ tags:
 - automatic-speech-recognition
 - mozilla-foundation/common_voice_8_0
 - generated_from_trainer
 datasets:
 - common_voice
 model-index:
-- name: ''
-  results: []
 ---
 <!-- This model card has been generated automatically according to the information the Trainer had access to. You
@@ -23,6 +39,8 @@ It achieves the following results on the evaluation set:
 - Loss: 0.5351
 - Wer: 2.6188
 ## Model description
 More information needed
@@ -73,3 +91,10 @@ The following hyperparameters were used during training:
 - Pytorch 1.10.2+cu102
 - Datasets 1.18.2.dev0
 - Tokenizers 0.11.0

 - automatic-speech-recognition
 - mozilla-foundation/common_voice_8_0
 - generated_from_trainer
+- robust-speech-event
+- ja
 datasets:
 - common_voice
 model-index:
+- name: 'XLS-R-300-m'
+  results:
+  - task:
+      name: Automatic Speech Recognition
+      type: automatic-speech-recognition
+    dataset:
+      name: Common Voice 8
+      type: mozilla-foundation/common_voice_8_0
+      args: ja
+    metrics:
+       - name: Test WER
+         type: wer
+         value: 94.91
+       - name: Test CER
+         type: cer
+         value: 23.32
 ---
 <!-- This model card has been generated automatically according to the information the Trainer had access to. You
 - Loss: 0.5351
 - Wer: 2.6188
+Kanji are converted into Hiragana using the [pykakasi](https://pykakasi.readthedocs.io/en/latest/index.html) library during training and evaluation. The model can output both Hiragana and Katakana characters.
 ## Model description
 More information needed
 - Pytorch 1.10.2+cu102
 - Datasets 1.18.2.dev0
 - Tokenizers 0.11.0
+#### Evaluation Commands
+1. To evaluate on `mozilla-foundation/common_voice_8_0` with split `test`
+```bash
+python ./eval.py --model_id AndrewMcDowell/wav2vec2-xls-r-300m-japanese --dataset mozilla-foundation/common_voice_8_0 --config ja --split test --log_outputs
+```