base-1e18-d896-seed46

16 effective layers, trained at a 1e18 FLOP budget at this architecture's compute-optimal width. One of a six-seed set (seeds 42-47) in which only the training data order varies; model initialisation is fixed across seeds.

field value
architecture base
d_model 896
d_ff 2432
width_ratio 7.0
base_d_model 128
base_d_ff 384
data seed 46
init seed 42

Loads with trust_remote_code=True. Trained with context length 1024; set max_length=1024 when evaluating.

Downloads last month
-
Safetensors
Model size
0.2B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support