YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
| Hyperparameter | Value |
|---|---|
| Epochs | 2 |
| Maximum sequence length | 16,384 |
| Loss | NLL, assistant tokens only |
| Sequence packing | Enabled |
| Per-device train batch size | 22 |
| Per-device evaluation batch size | 8 |
| Training processes | 8 |
| Gradient accumulation steps | 1 |
| Effective global train batch size | 176 |
| Optimizer | AdamW, fused Torch implementation |
| Learning rate | 1e-5 |
| Minimum learning rate | 1e-6 |
| Scheduler | Cosine warmup with minimum learning rate |
| Warmup ratio | 0.05 |
| Weight decay | 0.05 |
| Adam beta 1 | 0.9 |
| Adam beta 2 | 0.999 |
| Adam epsilon | 1e-8 |
| Maximum gradient norm | 1.0 |
| Precision | bfloat16 |
| Gradient checkpointing | Enabled, non-reentrant |
| Attention implementation | FlashAttention 2 |
| Liger kernels | Enabled |
| Distributed strategy | DeepSpeed ZeRO-2, no offload |
| Shuffle dataset | Enabled |
| Seed | 42 |
| Evaluation interval | 250 steps |
| Logging interval | 1 step |
| Save strategy | Every epoch |
| Save limit | 10 checkpoints |
| Dataset preprocessing workers | 80 |
- Downloads last month
- 33
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support