mlpr-qwen2.5-0.5b-instruct-50ep-adaptive

This model is a fine-tuned version of Qwen/Qwen2.5-0.5B-Instruct on the None dataset. It achieves the following results on the evaluation set:

  • Loss: 4.7677

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 2e-05
  • train_batch_size: 4
  • eval_batch_size: 4
  • seed: 42
  • gradient_accumulation_steps: 4
  • total_train_batch_size: 16
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: linear
  • lr_scheduler_warmup_ratio: 0.1
  • num_epochs: 50

Training results

Training Loss Epoch Step Validation Loss
4.5989 1.0 75 5.0353
1.8474 2.0 150 3.3853
0.6564 3.0 225 3.4776
0.6084 4.0 300 3.8776
0.6236 5.0 375 3.9364
0.5925 6.0 450 3.9675
0.61 7.0 525 4.0106
0.6006 8.0 600 3.9648
0.5914 9.0 675 4.0257
0.603 10.0 750 4.1859
0.5969 11.0 825 4.1352
0.5963 12.0 900 4.1591
0.6133 13.0 975 4.0093
0.5833 14.0 1050 4.0985
0.5841 15.0 1125 4.1116
0.5882 16.0 1200 4.5443
0.6036 17.0 1275 4.3679
0.5762 18.0 1350 4.6143
0.5814 19.0 1425 4.4029
0.5378 20.0 1500 4.3798
0.5598 21.0 1575 4.5965
0.498 22.0 1650 4.4821
0.4918 23.0 1725 4.5936
0.4452 24.0 1800 4.5116
0.4085 25.0 1875 4.4599
0.3966 26.0 1950 4.5074
0.37 27.0 2025 4.4912
0.3565 28.0 2100 4.5886
0.3757 29.0 2175 4.7237
0.3304 30.0 2250 4.6918
0.333 31.0 2325 4.6831
0.3221 32.0 2400 4.6908
0.2984 33.0 2475 4.6874
0.3155 34.0 2550 4.7677

Framework versions

  • PEFT 0.12.0
  • Transformers 4.44.2
  • Pytorch 2.5.1+cu124
  • Datasets 2.21.0
  • Tokenizers 0.19.1
Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for codewithdark/mlpr-qwen2.5-0.5b-instruct-50ep-adaptive

Adapter
(770)
this model