YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Celerity 271M 8K β€” LR ablation β€” ad0.2_lr1.2

Official Celerity 271M / 8K learning-rate-ablation checkpoint converted from Cerebras CS format to Hugging Face format.

Training provenance

  • Original experiment: ad0.2_lr1.2
  • Source checkpoint: checkpoint_13773.mdl
  • Attention dropout rate: 0.2
  • Attention-dropout schedule: constant
  • Peak learning rate: 1.2
  • Learning-rate multiplier relative to 0.15: 8
  • Weight decay: 0.00034673267902102944
  • tau_ema: 0.1745
  • Global train batch size: 48
  • Validation batch size: 32
  • Training steps: 13773
  • Maximum sequence length: 8192
  • Position embedding type: ALiBi
  • Residual dropout rate: 0.0
  • Stochastic depth: 0.0
  • LayerDrop: 0.0
  • Source runtime experiments: cbcore 2.6.0
  • Converter commit: 6f36ba6d76a97383171df68af096670b64f718b7

This LR ablation intentionally adjusts weight decay inversely with learning rate so that tau_ema remains fixed at 0.1745.

Attention dropout, learning rate, and weight decay are training-time hyperparameters. Their effects are represented in the learned checkpoint weights. Hugging Face evaluation is performed with dropout disabled by model.eval().

The model uses custom Celerity Hugging Face modeling code and should be loaded with trust_remote_code=True.

Downloads last month
18
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support