A-glu-linear-46L

This model is a fine-tuned version of on an unknown dataset. It achieves the following results on the evaluation set:

  • Loss: 2.0422

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 0.001
  • train_batch_size: 256
  • eval_batch_size: 256
  • seed: 42
  • gradient_accumulation_steps: 2
  • total_train_batch_size: 512
  • optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: constant
  • training_steps: 1000

Training results

Training Loss Epoch Step Validation Loss
12.1645 0.0270 50 5.8479
10.5654 0.0539 100 5.0850
9.3252 0.0809 150 4.4189
8.0606 0.1078 200 3.9415
7.5172 0.1348 250 3.6139
6.7608 0.1618 300 3.3349
6.3626 0.1887 350 3.0817
5.8267 0.2157 400 2.8769
5.5577 0.2427 450 2.7150
5.1979 0.2696 500 2.5881
5.0257 0.2966 550 2.4825
4.8073 0.3235 600 2.3988
4.6867 0.3505 650 2.3180
4.5310 0.3775 700 2.2552
4.4490 0.4044 750 2.2072
4.3345 0.4314 800 2.1641
4.2768 0.4583 850 2.1301
4.1870 0.4853 900 2.0952
4.1466 0.5123 950 2.0662
4.0888 0.5392 1000 2.0422

Framework versions

  • Transformers 5.15.0.dev0
  • Pytorch 2.6.0+cu124
  • Datasets 5.0.1
  • Tokenizers 0.22.2
Downloads last month
-
Safetensors
Model size
8.07M params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support