rostlabs/rost-286m-base

286M parameter nanochat GPT (base stage), research tag normalized-d12, checkpoint step 001920.

Part of the rost model zoo โ€” the full set of trained variants behind the rost research series, published for reproducibility. This is a research model, not a product.

architecture nanochat GPT, 12 layers, 768 embed, 2048 context
tokenizer included under tokenizer/ (32,768 vocab)
training mixture Romanian (FineWeb2-ro, educational score >= 3, diacritic-normalized)
stage base
val bpb (own split) 0.84148

The strongest base model of the rost small series (research tag normalized-d12). Cross-evaluated on held-out Romanian it outperforms its diacritic-stripped twin rost-286m-diacritics-stripped by ~60% bits-per-byte โ€” the headline result of the rost diacritics study.

Validation bits-per-byte is measured on this arm's own validation split and is not comparable across arms โ€” cross-arm comparisons in the rost write-ups are always cross-evaluated on identical text.

Licence

CC-BY-NC-4.0, non-commercial, inherited from the most restrictive component of the training data.

Downloads last month
7
Inference Providers NEW
This model isn't deployed by any Inference Provider. ๐Ÿ™‹ Ask for provider support