wav2vec2-xlsr-hre-v1

This model is a fine-tuned version of facebook/wav2vec2-large-xlsr-53 on the hre-audio-dataset8 dataset. It achieves the following results on the evaluation set:

  • Loss: 0.9722
  • Cer Ortho: 57.1608
  • Cer: 48.9420

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 0.0003
  • train_batch_size: 16
  • eval_batch_size: 16
  • seed: 42
  • optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • lr_scheduler_warmup_steps: 100
  • num_epochs: 20
  • mixed_precision_training: Native AMP

Training results

Training Loss Epoch Step Validation Loss Cer Ortho Cer
0.3674 0.4608 100 0.9867 61.0553 54.7010
0.3941 0.9217 200 0.9548 63.8729 54.5722
0.3239 1.3825 300 0.9726 60.3733 53.8178
0.4434 1.8433 400 0.8731 59.9605 53.5787
0.3304 2.3041 500 1.0734 64.0883 54.7930
0.3552 2.7650 600 0.8890 62.6525 53.1555
0.3390 3.2258 700 0.9412 58.5786 53.7994
0.3741 3.6866 800 0.9152 56.8737 52.6035
0.2740 4.1475 900 0.9503 57.8428 52.1803
0.2942 4.6083 1000 0.9740 57.8607 52.2355
0.3040 5.0691 1100 0.8920 58.9555 51.3707
0.3025 5.5300 1200 0.9718 58.5068 52.1067
0.2609 5.9908 1300 0.9868 59.3144 52.2171
0.2534 6.4516 1400 1.0018 56.7839 52.4011
0.3303 6.9124 1500 0.9764 58.3632 51.3707
0.2839 7.3733 1600 0.9399 60.0503 51.4811
0.2668 7.8341 1700 0.9931 60.9476 52.2355
0.3022 8.2949 1800 0.9181 60.4810 51.3891
0.3641 8.7558 1900 0.8288 58.0043 50.4876
0.3030 9.2166 2000 0.9127 58.8299 51.2052
0.3124 9.6774 2100 0.9390 56.6762 51.3707
0.2614 10.1382 2200 0.9390 58.8658 51.3339
0.2920 10.5991 2300 0.8887 58.0402 50.0644
0.2787 11.0599 2400 0.9710 61.3783 50.8372
0.2660 11.5207 2500 1.0166 62.5090 51.8675
0.2955 11.9816 2600 0.9085 60.4630 50.3404
0.2702 12.4424 2700 0.9840 60.2656 50.2300
0.2810 12.9032 2800 0.9933 59.5477 50.0092
0.2654 13.3641 2900 0.9433 59.5298 50.1932
0.2132 13.8249 3000 0.9613 57.7710 49.6964
0.1992 14.2857 3100 0.9687 58.7940 49.6596
0.2426 14.7465 3200 0.9831 59.8169 50.6716
0.2276 15.2074 3300 0.9570 58.5248 49.3284
0.1987 15.6682 3400 1.0591 58.5068 50.4692
0.2157 16.1290 3500 1.0039 57.9505 50.1012
0.1643 16.5899 3600 0.9512 56.8378 48.9604
0.2012 17.0507 3700 0.9912 57.3941 49.6412
0.1917 17.5115 3800 1.0092 57.5915 49.5860
0.1653 17.9724 3900 0.9984 57.5556 49.4572
0.1890 18.4332 4000 0.9601 57.1967 49.1812
0.1715 18.8940 4100 1.0008 57.5556 49.5676
0.1785 19.3548 4200 0.9669 57.0172 48.9052
0.1341 19.8157 4300 0.9722 57.2146 48.9604
0.1606 20.0 4340 0.9722 57.1608 48.9420

Framework versions

  • Transformers 5.13.1
  • Pytorch 2.11.0+cu128
  • Datasets 2.18.0
  • Tokenizers 0.22.2
Downloads last month
12
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ntviet/wav2vec2-xlsr-hre-v1

Finetuned
(391)
this model