byT5-large English→Dhivehi (sentence-level)

byT5-large (1.2B, byte-level encoder-decoder) for English→Dhivehi translation, trained on a sentence-level Dhivehi–English corpus (machine-translated).

Byte-level single-sentence model — do not feed multi-sentence input (it collapses on paragraphs). For multi-sentence input, see the paragraph models: Qwen-para / mT5-para. Single-sentence input only.

Scores (chrF / chrF++ / BLEU)

Benchmark chrF chrF++ BLEU
gold (human references, article-level, N=500) 46.89 38.89 3.25
held-out chunk (in-distribution) 67.74 60.74 21.34
held-out sentence (in-distribution) 64.41 57.76 19.91

chrF is the metric to trust for Thaana; BLEU is unreliable (word segmentation / morphology).

Example

Input (en): While this is a 7.6 percent increase compared to the same period last year, an average of 7,778 tourists visit the Maldives daily.

Output (dv): މިއީ މިދިޔަ އަހަރުގެ މި މުއްދަތާ އަޅާބަލާއިރު 7.6 އިންސައްތައިގެ ކުރިއެރުމެއް ކަމަށްވާއިރު، ދުވާލަކު އެވްރެޖްކޮށް 7،778 ފަތުރުވެރިން ރާއްޖެ ޒިޔާރަތްކުރެއެވެ.

Real held-out sample and this model's own output.

Usage

import torch
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

m = "Neobe/en-dhivehi-byt5-large-sentence"
tok = AutoTokenizer.from_pretrained(m)
model = AutoModelForSeq2SeqLM.from_pretrained(m, torch_dtype=torch.float32).eval().cuda()  # fp32

src = "The President of the Maldives met with the cabinet today."
inp = tok(src, return_tensors="pt", truncation=True, max_length=1024).to("cuda")
out = model.generate(**inp, max_new_tokens=512, num_beams=4, length_penalty=1.0, no_repeat_ngram_size=0)
print(tok.decode(out[0], skip_special_tokens=True))

Training

Base google/byt5-large; fp32; Adafactor; LR 1e-4 cosine; max_length 1024; 1 epoch; effective batch ~32; gradient checkpointing.

Limitations

Domain = Maldivian news / press / Wikipedia; technical or informal English is out of distribution. Non-human references are machine-generated (distillation). This byte-level model is single-sentence only.

Citation

@misc{neobe_en_dhivehi_byt5_large_sentence_2026,
  title  = {byT5-large English→Dhivehi (sentence-level)},
  author = {Neobe},
  year   = {2026},
  howpublished = {\url{https://huggingface.co/Neobe/en-dhivehi-byt5-large-sentence}}
}
Downloads last month
12
Safetensors
Model size
1B params
Tensor type
BF16
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Neobe/en-dhivehi-byt5-large-sentence

Finetuned
(31)
this model

Collection including Neobe/en-dhivehi-byt5-large-sentence