A-glu-waleed-9L

This model is a fine-tuned version of on an unknown dataset. It achieves the following results on the evaluation set:

  • Loss: 2.2273

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 0.001
  • train_batch_size: 128
  • eval_batch_size: 128
  • seed: 42
  • gradient_accumulation_steps: 4
  • total_train_batch_size: 512
  • optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: constant
  • training_steps: 1000

Training results

Training Loss Epoch Step Validation Loss
23.5068 0.0270 50 5.2269
16.9949 0.0539 100 4.1003
15.0230 0.0809 150 3.5819
13.3478 0.1078 200 3.2912
12.6337 0.1348 250 3.0792
11.7513 0.1618 300 2.9125
11.3232 0.1887 350 2.7897
10.8105 0.2157 400 2.6903
10.5763 0.2427 450 2.6156
10.2123 0.2696 500 2.5474
10.0600 0.2966 550 2.4986
9.8177 0.3235 600 2.4492
9.6829 0.3505 650 2.4087
9.5042 0.3775 700 2.3720
9.4019 0.4044 750 2.3435
9.2354 0.4314 800 2.3132
9.1831 0.4583 850 2.2884
9.0467 0.4853 900 2.2660
9.0000 0.5123 950 2.2443
8.9161 0.5392 1000 2.2273

Framework versions

  • Transformers 5.15.0.dev0
  • Pytorch 2.6.0+cu124
  • Datasets 5.0.1
  • Tokenizers 0.22.2
Downloads last month
156
Safetensors
Model size
2M params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support