base-1e18-d896-seed45
16 effective layers, trained at a 1e18 FLOP budget at this architecture's compute-optimal width. One of a six-seed set (seeds 42-47) in which only the training data order varies; model initialisation is fixed across seeds.
| field | value |
|---|---|
| architecture | base |
| d_model | 896 |
| d_ff | 2432 |
| width_ratio | 7.0 |
| base_d_model | 128 |
| base_d_ff | 384 |
| data seed | 45 |
| init seed | 42 |
Loads with trust_remote_code=True. Trained with context length 1024; set
max_length=1024 when evaluating.
- Downloads last month
- 39
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support