rostlabs/rost-73m-diacritics-normalized

73M parameter nanochat GPT (base stage), research tag normalized-d6, checkpoint step 001920.

Part of the rost model zoo — the full set of trained variants behind the rost research series, published for reproducibility. This is a research model, not a product.

architecture nanochat GPT, 6 layers, 384 embed, 2048 context
tokenizer included under tokenizer/ (32,768 vocab)
training mixture Romanian, cedilla forms normalized to comma-below (sÈ™/È› kept)
stage base
val bpb (own split) 1.03736

Diacritics-study arm (research tag normalized-d6), scale sweep point at 73M. Pair with rost-73m-diacritics-stripped: the stripped twin pays ~54% more bits-per-byte on correct Romanian at this scale.

Validation bits-per-byte is measured on this arm's own validation split and is not comparable across arms — cross-arm comparisons in the rost write-ups are always cross-evaluated on identical text.

Licence

CC-BY-NC-4.0, non-commercial, inherited from the most restrictive component of the training data.

Downloads last month
-
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support