small

This model is a fine-tuned version of meta-llama/Llama-3.1-8B-Instruct on the small_data_new_train dataset. It achieves the following results on the evaluation set:

  • Loss: 1.4002

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 0.0001
  • train_batch_size: 1
  • eval_batch_size: 1
  • seed: 42
  • distributed_type: multi-GPU
  • gradient_accumulation_steps: 8
  • total_train_batch_size: 8
  • optimizer: Use OptimizerNames.ADAMW_TORCH with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: cosine
  • lr_scheduler_warmup_ratio: 0.1
  • num_epochs: 100

Training results

Training Loss Epoch Step Validation Loss
No log 0.4 1 2.9306
No log 0.8 2 2.9292
No log 1.0 3 2.9292
No log 1.4 4 2.9132
No log 1.8 5 2.9042
No log 2.0 6 2.8854
No log 2.4 7 2.8500
No log 2.8 8 2.7950
No log 3.0 9 2.7950
2.7586 3.4 10 2.7017
2.7586 3.8 11 2.5546
2.7586 4.0 12 2.3694
2.7586 4.4 13 2.1354
2.7586 4.8 14 1.8584
2.7586 5.0 15 1.8584
2.7586 5.4 16 1.6423
2.7586 5.8 17 1.5388
2.7586 6.0 18 1.4914
2.7586 6.4 19 1.4529
1.3615 6.8 20 1.4321
1.3615 7.0 21 1.4321
1.3615 7.4 22 1.4145
1.3615 7.8 23 1.4052
1.3615 8.0 24 1.3982
1.3615 8.4 25 1.4006
1.3615 8.8 26 1.3924
1.3615 9.0 27 1.3924
1.3615 9.4 28 1.4008
1.3615 9.8 29 1.4002

Framework versions

  • PEFT 0.15.1
  • Transformers 4.51.1
  • Pytorch 2.7.1+cu126
  • Datasets 3.5.0
  • Tokenizers 0.21.0
Downloads last month
6
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Gnomy19/llama3_small_data

Adapter
(2917)
this model