nllb_ipa_to_bangla_epoch25

This model is a fine-tuned version of facebook/nllb-200-distilled-600M on the None dataset. It achieves the following results on the evaluation set:

  • Loss: 2.1860
  • Bleu: 15.0397
  • Chrf: 40.4001
  • Meteor: 0.3216
  • Bertscore F1: 0.8476

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 5e-05
  • train_batch_size: 8
  • eval_batch_size: 8
  • seed: 42
  • optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • lr_scheduler_warmup_steps: 50
  • num_epochs: 25
  • mixed_precision_training: Native AMP

Training results

Training Loss Epoch Step Validation Loss Bleu Chrf Meteor Bertscore F1
2.9664 1.0 189 2.4330 1.1567 17.0736 0.0909 0.7679
1.9794 2.0 378 2.1012 4.5012 23.9044 0.1731 0.795
1.4515 3.0 567 2.0078 6.1795 28.9153 0.2162 0.8106
1.1065 4.0 756 1.9566 10.5487 32.5721 0.2395 0.8193
0.8437 5.0 945 1.9606 9.2165 32.1324 0.2607 0.8245
0.6498 6.0 1134 1.9997 11.145 34.7027 0.2594 0.8293
0.4911 7.0 1323 2.0180 12.3754 36.4547 0.2842 0.8327
0.3768 8.0 1512 2.0067 12.4325 37.4606 0.2817 0.8358
0.2899 9.0 1701 2.0370 15.0857 38.4527 0.3063 0.8408
0.2198 10.0 1890 2.0621 13.5259 37.2397 0.2933 0.8366
0.1704 11.0 2079 2.1010 14.3345 37.9934 0.2929 0.8391
0.1355 12.0 2268 2.1056 14.9335 38.6265 0.3077 0.8418
0.1118 13.0 2457 2.1115 13.5777 38.8908 0.3181 0.8425
0.0906 14.0 2646 2.1289 14.5873 38.5979 0.3099 0.8431
0.0748 15.0 2835 2.1332 14.1024 39.7978 0.3117 0.8422
0.0638 16.0 3024 2.1610 15.3784 39.4667 0.3131 0.8422
0.0571 17.0 3213 2.1501 15.7275 40.414 0.3245 0.8476
0.0519 18.0 3402 2.1482 16.9568 40.3169 0.3142 0.848
0.0468 19.0 3591 2.1651 15.4548 40.654 0.323 0.847
0.0399 20.0 3780 2.1768 15.4357 40.1039 0.3233 0.8499
0.0378 21.0 3969 2.1720 14.9499 40.049 0.3148 0.8481
0.0356 22.0 4158 2.1779 15.3876 40.6863 0.3256 0.8499
0.0329 23.0 4347 2.1900 15.6909 41.2235 0.3275 0.8497
0.0309 24.0 4536 2.1893 14.8779 40.3529 0.319 0.847
0.0300 25.0 4725 2.1860 15.0397 40.4001 0.3216 0.8476

Framework versions

  • Transformers 5.15.0
  • Pytorch 2.11.0+cu128
  • Datasets 4.0.0
  • Tokenizers 0.22.2
Downloads last month
122
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for thunderboltc/nllb_ipa_to_bangla_epoch25

Finetuned
(372)
this model