English → Rohingya (opus-mt finetune)

Translates English into Rohingya. Finetuned by Linguardia from Helsinki-NLP/opus-mt-en-mul on the English↔Rohingya sentence pairs in Tatoeba.

Built to generate course material for a language that no major machine translator supports.

Usage

from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

tok = AutoTokenizer.from_pretrained("linguardia/opus-mt-en-rhg")
model = AutoModelForSeq2SeqLM.from_pretrained("linguardia/opus-mt-en-rhg")

batch = tok([">>rhg<< I wake up early."], return_tensors="pt")
out = model.generate(**batch, num_beams=4, max_length=64)
print(tok.batch_decode(out, skip_special_tokens=True))

The target language is selected by the >>rhg<< prefix on the source text. It was added as a new token rather than replacing one of the base model's existing languages, so every language the base already handled still works.

Training

base Helsinki-NLP/opus-mt-en-mul (77M params)
training pairs 3,548
held out 300 dev, 0 gold
epochs 20.0
batch size 32
optimiser Adafactor, lr 5e-5, 500 warmup steps
max sequence 48 tokens (99.9% of source, 99.6% of target)
hardware Apple M4, MPS

Evaluation

Held-out Tatoeba (300 pairs never seen in training):

metric score
BLEU 18.4
chrF 42.5

Translating the same 2,000-sentence course corpus:

corpus rows distinct vocab words/sentence
this model 2000 1998 (99.9%) 2246 6.7

Training data and attribution

Sentence pairs come from Tatoeba, released under CC-BY 2.0 FR, with some sentences under CC0 1.0. Tatoeba's sentences are written by volunteers, and this model is a derivative of their work: if you use it, credit Tatoeba and its contributors.

Pairs were deduplicated on exact (English, Rohingya) text. No machine translation was used to create the training data at any point.

Licence chain

layer licence
base model Helsinki-NLP/opus-mt-en-mul Apache-2.0
training data (Tatoeba) CC-BY 2.0 FR / CC0 1.0
these weights Apache-2.0, with the attribution above

Limitations

  • Trained on Tatoeba, which is short everyday sentences. Expect it to be weakest on long, technical or literary text.
  • Neural MT translates each sentence independently, so a term may be rendered differently across a corpus. Check consistency if that matters to you.
  • Output should be reviewed by a speaker before being taught to learners.
Downloads last month
5
Safetensors
Model size
77M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for linguardia/opus-mt-en-rhg

Finetuned
(25)
this model

Dataset used to train linguardia/opus-mt-en-rhg