QBert

This model is a fine-tuned version of on an unknown dataset. It achieves the following results on the evaluation set:

  • Loss: 0.4291

Model description

More information needed

Intended uses & limitations

More information needed

Training and evaluation data

More information needed

Training procedure

Training hyperparameters

The following hyperparameters were used during training:

  • learning_rate: 0.0001
  • train_batch_size: 64
  • eval_batch_size: 8
  • seed: 42
  • optimizer: Use OptimizerNames.ADAMW_TORCH_FUSED with betas=(0.9,0.999) and epsilon=1e-08 and optimizer_args=No additional optimizer arguments
  • lr_scheduler_type: linear
  • lr_scheduler_warmup_steps: 10000
  • num_epochs: 603
  • mixed_precision_training: Native AMP

Training results

Training Loss Epoch Step Validation Loss
6.7876 4.0161 2000 5.7572
5.6570 8.0321 4000 5.6031
5.4869 12.0482 6000 5.0994
4.7571 16.0643 8000 4.4091
4.2246 20.0803 10000 3.9246
3.6582 24.0964 12000 3.1099
2.6372 28.1124 14000 2.0094
1.8545 32.1285 16000 1.5949
1.4873 36.1446 18000 1.3468
1.2508 40.1606 20000 1.1876
1.0889 44.1767 22000 1.0385
0.9580 48.1928 24000 0.9333
0.8578 52.2088 26000 0.8454
0.7791 56.2249 28000 0.7766
0.7070 60.2410 30000 0.7423
0.6498 64.2570 32000 0.6968
0.6035 68.2731 34000 0.6754
0.5591 72.2892 36000 0.6412
0.5222 76.3052 38000 0.5862
0.4908 80.3213 40000 0.5847
0.4632 84.3373 42000 0.5529
0.4383 88.3534 44000 0.5341
0.4163 92.3695 46000 0.5323
0.3944 96.3855 48000 0.5072
0.3762 100.4016 50000 0.4913
0.3596 104.4177 52000 0.4944
0.3429 108.4337 54000 0.4898
0.3278 112.4498 56000 0.4626
0.3164 116.4659 58000 0.4516
0.3056 120.4819 60000 0.4548
0.2926 124.4980 62000 0.4538
0.2825 128.5141 64000 0.4484
0.2747 132.5301 66000 0.4448
0.2643 136.5462 68000 0.4394
0.2554 140.5622 70000 0.4172
0.2484 144.5783 72000 0.4299
0.2422 148.5944 74000 0.4298
0.2345 152.6104 76000 0.4291

Framework versions

  • Transformers 5.6.2
  • Pytorch 2.11.0+cu130
  • Datasets 4.8.4
  • Tokenizers 0.22.2
Downloads last month
6
Safetensors
Model size
68M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support