English → Tigre (opus-mt finetune)

Translates English into Tigre. Finetuned by Linguardia from Helsinki-NLP/opus-mt-en-mul on the English↔Tigre sentence pairs in Tatoeba.

Built to generate course material for a language that no major machine translator supports.

Usage

from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

tok = AutoTokenizer.from_pretrained("linguardia/opus-mt-en-tig")
model = AutoModelForSeq2SeqLM.from_pretrained("linguardia/opus-mt-en-tig")

batch = tok([">>tig<< I wake up early."], return_tensors="pt")
out = model.generate(**batch, num_beams=4, max_length=64)
print(tok.batch_decode(out, skip_special_tokens=True))

The target language is selected by the >>tig<< prefix on the source text. It was added as a new token rather than replacing one of the base model's existing languages, so every language the base already handled still works.

Training

base Helsinki-NLP/opus-mt-en-mul (77M params)
training pairs 10,006
held out 526 dev, 9 gold
epochs 8.0
batch size 32
optimiser Adafactor, lr 5e-5, 500 warmup steps
max sequence 48 tokens (99.9% of source, 99.6% of target)
hardware Apple M4, MPS

Evaluation

Held-out Tatoeba (526 pairs never seen in training):

metric score
BLEU 20.3
chrF 35.1

Against 9 human translations of sentences that appear verbatim in the target corpus:

metric score
BLEU 10.6
chrF 18.0

Translating the same 2,000-sentence course corpus:

corpus rows distinct vocab words/sentence
this model 2000 1999 (100.0%) 2853 5.1

Training data and attribution

Sentence pairs come from Tatoeba, released under CC-BY 2.0 FR, with some sentences under CC0 1.0. Tatoeba's sentences are written by volunteers, and this model is a derivative of their work: if you use it, credit Tatoeba and its contributors.

Pairs were deduplicated on exact (English, Tigre) text. No machine translation was used to create the training data at any point.

Licence chain

layer licence
base model Helsinki-NLP/opus-mt-en-mul Apache-2.0
training data (Tatoeba) CC-BY 2.0 FR / CC0 1.0
these weights Apache-2.0, with the attribution above

Limitations

  • Trained on Tatoeba, which is short everyday sentences. Expect it to be weakest on long, technical or literary text.
  • Neural MT translates each sentence independently, so a term may be rendered differently across a corpus. Check consistency if that matters to you.
  • Output should be reviewed by a speaker before being taught to learners.
Downloads last month
-
Safetensors
Model size
77M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for linguardia/opus-mt-en-tig

Finetuned
(25)
this model

Dataset used to train linguardia/opus-mt-en-tig