ec9a614c892a5ce57e8eb6439ad3bd51

This model is a fine-tuned version of google/umt5-small on the Helsinki-NLP/opus_books [fr-pt] dataset. It achieves the following results on the evaluation set:

  • Loss: 2.8520
  • Data Size: 1.0
  • Epoch Runtime: 8.0758
  • Bleu: 5.1945

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 5e-05
  • train_batch_size: 8
  • eval_batch_size: 8
  • seed: 42
  • distributed_type: multi-GPU
  • num_devices: 4
  • total_train_batch_size: 32
  • total_eval_batch_size: 32
  • optimizer: Use adamw_torch with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: constant
  • num_epochs: 50

Training results

Training Loss Epoch Step Validation Loss Data Size Epoch Runtime Bleu
No log 0 0 16.1748 0 1.1777 0.4226
No log 1 31 16.1153 0.0078 1.3125 0.4585
No log 2 62 15.9533 0.0156 1.6327 0.4890
No log 3 93 15.8908 0.0312 1.9235 0.5583
No log 4 124 15.8601 0.0625 2.0207 0.3142
No log 5 155 15.5561 0.125 2.6232 0.3385
No log 6 186 15.6011 0.25 3.4505 0.3442
2.9313 7 217 14.4831 0.5 4.8066 0.3555
2.9313 8.0 248 11.6977 1.0 7.6944 0.3809
10.718 9.0 279 9.6328 1.0 7.8142 0.3768
12.0868 10.0 310 8.1565 1.0 8.3106 0.3889
12.0868 11.0 341 7.1047 1.0 8.9109 0.2825
9.558 12.0 372 6.3523 1.0 8.9924 0.4299
8.0857 13.0 403 5.4300 1.0 9.0836 0.9043
8.0857 14.0 434 4.7490 1.0 6.3490 2.0336
7.0196 15.0 465 4.4147 1.0 6.6850 3.5048
7.0196 16.0 496 4.2541 1.0 6.8621 4.3999
6.2394 17.0 527 4.1485 1.0 7.2159 4.8407
5.7914 18.0 558 4.0515 1.0 8.4663 5.1538
5.7914 19.0 589 3.9218 1.0 7.5790 5.6354
5.4274 20.0 620 3.8342 1.0 7.6054 3.8640
5.1882 21.0 651 3.7428 1.0 7.5607 2.1286
5.1882 22.0 682 3.6675 1.0 7.4714 1.7915
4.9689 23.0 713 3.5918 1.0 7.5897 1.8241
4.9689 24.0 744 3.5217 1.0 7.9090 1.8391
4.7617 25.0 775 3.4577 1.0 8.3472 1.2336
4.6285 26.0 806 3.3979 1.0 8.1950 1.2248
4.6285 27.0 837 3.3432 1.0 8.5740 1.2389
4.4538 28.0 868 3.2944 1.0 6.4040 1.3340
4.4538 29.0 899 3.2475 1.0 7.1924 1.3961
4.3293 30.0 930 3.2026 1.0 7.1671 1.4277
4.2152 31.0 961 3.1681 1.0 7.2722 1.7774
4.2152 32.0 992 3.1311 1.0 7.1802 3.3872
4.1074 33.0 1023 3.1071 1.0 7.5470 6.7424
4.0027 34.0 1054 3.0775 1.0 7.5095 5.6128
4.0027 35.0 1085 3.0529 1.0 7.5647 4.8327
3.9314 36.0 1116 3.0317 1.0 7.8623 4.6476
3.9314 37.0 1147 3.0122 1.0 7.8229 4.6880
3.8486 38.0 1178 3.0019 1.0 7.6763 4.6790
3.7769 39.0 1209 2.9852 1.0 8.2100 4.7554
3.7769 40.0 1240 2.9691 1.0 9.0335 4.8626
3.7066 41.0 1271 2.9450 1.0 8.9545 5.0563
3.6342 42.0 1302 2.9380 1.0 6.8437 5.0005
3.6342 43.0 1333 2.9260 1.0 6.9673 5.0035
3.5889 44.0 1364 2.9099 1.0 6.9993 5.0697
3.5889 45.0 1395 2.8987 1.0 7.4151 4.9567
3.5232 46.0 1426 2.8914 1.0 7.8934 4.9850
3.4865 47.0 1457 2.8773 1.0 8.4549 5.0317
3.4865 48.0 1488 2.8704 1.0 8.1827 5.1013
3.4057 49.0 1519 2.8602 1.0 7.9880 5.1254
3.3826 50.0 1550 2.8520 1.0 8.0758 5.1945

Framework versions

  • Transformers 4.57.0
  • Pytorch 2.8.0+cu128
  • Datasets 4.2.0
  • Tokenizers 0.22.1
Downloads last month
4
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for contemmcm/ec9a614c892a5ce57e8eb6439ad3bd51

Finetuned
(45)
this model