Are you sure there's enough credits?
20T is a lot of tokens
There isn't even enough credits currently to do it
ha ha no its not 20T
i made a naming mistake sorry
How much then?
its just gonna be 1-2T 20T is too much.
which context
it's in pretokenized chunks
2048 tokens.
its a bit short but i train in phases usally
in 2 days you could only:
8.6B tokens (8.8B with causal-aware attention).
Budget: 2 days × 3.2e15 FLOP/s ≈ 5.5e20 FLOPs, at ~6.5e10 FLOPs/token.
That's ~0.86 tokens/param — far below Chinchilla (20×). For a 2-day run on that node, a ~400M model at 8B tokens is much better matched.
no no! im not planning to use all. i only use a subset im well aware of the time constraints
hmm true
how many then?
i usually get data first so i dont have to worry about it later
If you have $316,800 then go and train on 1.4T tokens for 1 year
i plan the model after the dataa im planning a 3-10B ish parameter model ill test my pipeline first, then train in monthly shards
well i dont lol
your models are quite capable for their size i checked some out previously
I'd really reccomend like a smaller model on more tokens, then you could actually compete with other models. Thats what we do.
your profile says you have over 1900 TFLOPS of compute. Do you own 8 x H200s or is it cloud?
ok thanks for advice
Sure!
