English → Kabyle (opus-mt finetune)

Translates English into Kabyle. Finetuned by Linguardia from Helsinki-NLP/opus-mt-en-mul on the English↔Kabyle sentence pairs in Tatoeba.

Built to generate course material for a language that no major machine translator supports.

Usage

from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

tok = AutoTokenizer.from_pretrained("linguardia/opus-mt-en-kab")
model = AutoModelForSeq2SeqLM.from_pretrained("linguardia/opus-mt-en-kab")

batch = tok([">>kab<< I wake up early."], return_tensors="pt")
out = model.generate(**batch, num_beams=4, max_length=64)
print(tok.batch_decode(out, skip_special_tokens=True))

The target language is selected by the >>kab<< prefix on the source text. It was added as a new token rather than replacing one of the base model's existing languages, so every language the base already handled still works.

Training

base Helsinki-NLP/opus-mt-en-mul (77M params)
training pairs 138,740
held out 1,000 dev, 137 gold
epochs 2.0
batch size 32
optimiser Adafactor, lr 5e-5, 500 warmup steps
max sequence 48 tokens (99.9% of source, 99.6% of target)
hardware Apple M4, MPS

Evaluation

Held-out Tatoeba (1000 pairs never seen in training):

metric score
BLEU 26.2
chrF 50.9

Against 137 human translations of sentences that appear verbatim in the target corpus:

metric score
BLEU 22.9
chrF 44.2

Translating the same 2,000-sentence course corpus:

corpus rows distinct vocab words/sentence
this model 2000 1999 (100.0%) 2207 6.9

Training data and attribution

Sentence pairs come from Tatoeba, released under CC-BY 2.0 FR, with some sentences under CC0 1.0. Tatoeba's sentences are written by volunteers, and this model is a derivative of their work: if you use it, credit Tatoeba and its contributors.

Pairs were deduplicated on exact (English, Kabyle) text. No machine translation was used to create the training data at any point.

Licence chain

layer licence
base model Helsinki-NLP/opus-mt-en-mul Apache-2.0
training data (Tatoeba) CC-BY 2.0 FR / CC0 1.0
these weights Apache-2.0, with the attribution above

Limitations

  • Trained on Tatoeba, which is short everyday sentences. Expect it to be weakest on long, technical or literary text.
  • Neural MT translates each sentence independently, so a term may be rendered differently across a corpus. Check consistency if that matters to you.
  • Output should be reviewed by a speaker before being taught to learners.
Downloads last month
-
Safetensors
Model size
77M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for linguardia/opus-mt-en-kab

Finetuned
(23)
this model

Dataset used to train linguardia/opus-mt-en-kab