my_t5_small_test

This model was trained from scratch on the None dataset. It achieves the following results on the evaluation set:

  • Loss: 0.0284
  • Bleu: 94.5112
  • Gen Len: 28.0025

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 2e-05
  • train_batch_size: 16
  • eval_batch_size: 16
  • seed: 42
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: linear
  • num_epochs: 50
  • mixed_precision_training: Native AMP

Training results

Training Loss Epoch Step Validation Loss Bleu Gen Len
No log 1.0 25 0.2332 84.2783 26.685
No log 2.0 50 0.1832 87.343 27.29
No log 3.0 75 0.1488 88.2143 27.48
No log 4.0 100 0.1164 89.9573 28.0175
No log 5.0 125 0.0972 89.9952 27.9375
No log 6.0 150 0.0839 90.1522 27.975
No log 7.0 175 0.0708 90.5573 27.9775
No log 8.0 200 0.0634 91.0811 27.975
No log 9.0 225 0.0567 92.014 27.98
No log 10.0 250 0.0506 92.5003 28.0
No log 11.0 275 0.0471 92.6721 28.005
No log 12.0 300 0.0435 93.1332 28.1125
No log 13.0 325 0.0411 93.0801 28.01
No log 14.0 350 0.0391 93.4516 27.9675
No log 15.0 375 0.0375 93.6638 28.005
No log 16.0 400 0.0364 93.5842 28.0175
No log 17.0 425 0.0352 93.9818 28.09
No log 18.0 450 0.0342 94.273 28.1575
No log 19.0 475 0.0336 93.9818 28.08
0.1533 20.0 500 0.0330 93.9818 28.01
0.1533 21.0 525 0.0323 93.9553 28.05
0.1533 22.0 550 0.0320 93.8493 27.955
0.1533 23.0 575 0.0317 93.8758 27.955
0.1533 24.0 600 0.0312 93.8493 27.955
0.1533 25.0 625 0.0311 93.8758 27.8825
0.1533 26.0 650 0.0307 93.9818 27.93
0.1533 27.0 675 0.0305 94.1407 27.9925
0.1533 28.0 700 0.0303 94.1407 27.9475
0.1533 29.0 725 0.0301 94.1407 27.9075
0.1533 30.0 750 0.0299 94.273 27.9525
0.1533 31.0 775 0.0296 94.3524 28.0125
0.1533 32.0 800 0.0296 94.0877 27.895
0.1533 33.0 825 0.0294 94.1671 27.9275
0.1533 34.0 850 0.0292 94.1936 27.945
0.1533 35.0 875 0.0290 94.326 28.0475
0.1533 36.0 900 0.0289 94.3789 28.06
0.1533 37.0 925 0.0288 94.3789 28.0225
0.1533 38.0 950 0.0287 94.3789 28.035
0.1533 39.0 975 0.0286 94.3789 28.06
0.0545 40.0 1000 0.0285 94.4318 28.02
0.0545 41.0 1025 0.0285 94.4318 27.99
0.0545 42.0 1050 0.0285 94.3524 27.9625
0.0545 43.0 1075 0.0285 94.4053 27.975
0.0545 44.0 1100 0.0284 94.4847 27.9875
0.0545 45.0 1125 0.0284 94.5112 27.9875
0.0545 46.0 1150 0.0284 94.4847 27.9875
0.0545 47.0 1175 0.0284 94.4318 27.9825
0.0545 48.0 1200 0.0284 94.564 27.995
0.0545 49.0 1225 0.0284 94.5112 28.0025
0.0545 50.0 1250 0.0284 94.5112 28.0025

Framework versions

  • Transformers 4.44.2
  • Pytorch 2.4.1+cu121
  • Datasets 3.0.1
  • Tokenizers 0.19.1
Downloads last month
6
Safetensors
Model size
60.5M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support