mlpr-qwen2.5-7b-instruct-50ep-adaptive

This model is a fine-tuned version of Qwen/Qwen2.5-7B-Instruct on the None dataset. It achieves the following results on the evaluation set:

  • Loss: 4.5491

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 2e-05
  • train_batch_size: 4
  • eval_batch_size: 4
  • seed: 42
  • gradient_accumulation_steps: 4
  • total_train_batch_size: 16
  • optimizer: Adam with betas=(0.9,0.999) and epsilon=1e-08
  • lr_scheduler_type: linear
  • lr_scheduler_warmup_ratio: 0.1
  • num_epochs: 50

Training results

Training Loss Epoch Step Validation Loss
4.8327 1.0 75 5.0948
1.8797 2.0 150 3.6595
0.6155 3.0 225 3.6171
0.6003 4.0 300 3.9781
0.6183 5.0 375 3.9947
0.5854 6.0 450 4.0084
0.6 7.0 525 4.1125
0.5958 8.0 600 4.0350
0.5896 9.0 675 4.1435
0.5851 10.0 750 4.1408
0.5381 11.0 825 4.1460
0.5215 12.0 900 4.2615
0.5007 13.0 975 4.3515
0.4325 14.0 1050 4.2930
0.3926 15.0 1125 4.2548
0.374 16.0 1200 4.2715
0.3655 17.0 1275 4.2320
0.3455 18.0 1350 4.2249
0.3361 19.0 1425 4.2865
0.3302 20.0 1500 4.3143
0.3429 21.0 1575 4.2390
0.3171 22.0 1650 4.3041
0.3307 23.0 1725 4.2728
0.3086 24.0 1800 4.3276
0.3343 25.0 1875 4.2762
0.3121 26.0 1950 4.2976
0.3023 27.0 2025 4.3453
0.2966 28.0 2100 4.3390
0.3043 29.0 2175 4.2946
0.3044 30.0 2250 4.3264
0.3179 31.0 2325 4.3047
0.3008 32.0 2400 4.3251
0.3001 33.0 2475 4.3614
0.3035 34.0 2550 4.3776
0.3031 35.0 2625 4.3700
0.2955 36.0 2700 4.3790
0.299 37.0 2775 4.3477
0.294 38.0 2850 4.3304
0.2847 39.0 2925 4.3056
0.2872 40.0 3000 4.3254
0.2883 41.0 3075 4.3488
0.287 42.0 3150 4.4074
0.2773 43.0 3225 4.3696
0.2772 44.0 3300 4.4172
0.2709 45.0 3375 4.4326
0.2772 46.0 3450 4.4413
0.2533 47.0 3525 4.4771
0.2695 48.0 3600 4.5241
0.25 49.0 3675 4.5364
0.2588 50.0 3750 4.5491

Framework versions

  • PEFT 0.12.0
  • Transformers 4.44.2
  • Pytorch 2.5.1+cu124
  • Datasets 2.21.0
  • Tokenizers 0.19.1
Downloads last month
226
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for codewithdark/mlpr-qwen2.5-7b-instruct-50ep-adaptive

Base model

Qwen/Qwen2.5-7B
Adapter
(2604)
this model