lamso-en-nllb

This model is a fine-tuned version of facebook/nllb-200-distilled-600M on the None dataset. It achieves the following results on the evaluation set:

  • Loss: 3.1536
  • Bleu: 5.4661

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 2e-05
  • train_batch_size: 4
  • eval_batch_size: 4
  • seed: 42
  • gradient_accumulation_steps: 4
  • total_train_batch_size: 16
  • optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • lr_scheduler_warmup_steps: 50
  • num_epochs: 50
  • mixed_precision_training: Native AMP

Training results

Training Loss Epoch Step Validation Loss Bleu
5.5569 1.0 40 4.6347 0.0949
4.583 2.0 80 3.8912 0.2283
4.0616 3.0 120 3.5883 0.2882
3.698 4.0 160 3.4173 0.6949
3.5286 5.0 200 3.2860 1.0174
3.3239 6.0 240 3.2013 1.0944
3.2044 7.0 280 3.1336 1.1160
3.0838 8.0 320 3.1032 1.5282
2.982 9.0 360 3.0618 2.1756
2.84 10.0 400 3.0452 2.2410
2.7283 11.0 440 3.0317 2.4307
2.7494 12.0 480 3.0313 2.6356
2.5419 13.0 520 3.0275 2.3977
2.6218 14.0 560 3.0123 2.8629
2.5425 15.0 600 3.0041 2.8604
2.461 16.0 640 3.0067 3.3949
2.4507 17.0 680 3.0097 3.1292
2.2282 18.0 720 3.0183 3.5313
2.4489 19.0 760 3.0126 3.5240
2.144 20.0 800 3.0304 3.7798
2.1395 21.0 840 3.0248 4.0427
2.066 22.0 880 3.0380 3.9976
2.1631 23.0 920 3.0704 4.0978
2.038 24.0 960 3.0579 4.1935
2.0262 25.0 1000 3.0652 4.1297
2.0747 26.0 1040 3.0674 4.4140
1.8927 27.0 1080 3.0789 4.5061
1.8739 28.0 1120 3.0898 4.4716
1.8864 29.0 1160 3.0870 4.7618
1.9223 30.0 1200 3.0986 4.6710
1.9868 31.0 1240 3.1137 4.9288
1.9936 32.0 1280 3.1169 4.6553
1.9889 33.0 1320 3.1330 4.5970
1.906 34.0 1360 3.1260 4.7602
1.8468 35.0 1400 3.1294 4.6674
1.8345 36.0 1440 3.1246 5.0111
1.8759 37.0 1480 3.1325 4.5826
1.688 38.0 1520 3.1386 4.8248
1.9076 39.0 1560 3.1418 5.0801
1.681 40.0 1600 3.1414 5.1977
1.6575 41.0 1640 3.1469 5.1851
1.7737 42.0 1680 3.1456 4.9597
1.7835 43.0 1720 3.1512 5.2460
1.796 44.0 1760 3.1401 5.2646
1.8359 45.0 1800 3.1444 5.2608
1.7384 46.0 1840 3.1464 5.3784
1.724 47.0 1880 3.1515 5.3768
1.7753 48.0 1920 3.1536 5.4661
1.5805 49.0 1960 3.1526 5.3439
1.7253 50.0 2000 3.1536 5.2808

Framework versions

  • Transformers 4.57.2
  • Pytorch 2.9.0+cu126
  • Datasets 4.0.0
  • Tokenizers 0.22.1
Downloads last month
5
Safetensors
Model size
0.6B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Nichoh/lamso-en-nllb

Finetuned
(380)
this model