English → Lojban (opus-mt finetune)

Translates English into Lojban. Finetuned by Linguardia from Helsinki-NLP/opus-mt-en-mul on the English↔Lojban sentence pairs in Tatoeba.

Built to generate course material for a language that no major machine translator supports.

Usage

from transformers import AutoTokenizer, AutoModelForSeq2SeqLM

tok = AutoTokenizer.from_pretrained("linguardia/opus-mt-en-jbo")
model = AutoModelForSeq2SeqLM.from_pretrained("linguardia/opus-mt-en-jbo")

batch = tok([">>jbo<< I wake up early."], return_tensors="pt")
out = model.generate(**batch, num_beams=4, max_length=64)
print(tok.batch_decode(out, skip_special_tokens=True))

The target language is selected by the >>jbo<< prefix on the source text. It was added as a new token rather than replacing one of the base model's existing languages, so every language the base already handled still works.

Training

base Helsinki-NLP/opus-mt-en-mul (77M params)
training pairs 12,620
held out 1,000 dev, 39 gold
epochs 6.0
batch size 32
optimiser Adafactor, lr 5e-5, 500 warmup steps
max sequence 48 tokens (99.9% of source, 99.6% of target)
hardware Apple M4, MPS

Evaluation

Held-out Tatoeba (1000 pairs never seen in training):

metric score
BLEU 24.9
chrF 47.4

Against 39 human translations of sentences that appear verbatim in the target corpus:

metric score
BLEU 35.7
chrF 52.1

Translating the same 2,000-sentence course corpus:

corpus rows distinct vocab words/sentence
this model 2000 1999 (100.0%) 1076 8.7

Training data and attribution

Sentence pairs come from Tatoeba, released under CC-BY 2.0 FR, with some sentences under CC0 1.0. Tatoeba's sentences are written by volunteers, and this model is a derivative of their work: if you use it, credit Tatoeba and its contributors.

Pairs were deduplicated on exact (English, Lojban) text. No machine translation was used to create the training data at any point.

Licence chain

layer licence
base model Helsinki-NLP/opus-mt-en-mul Apache-2.0
training data (Tatoeba) CC-BY 2.0 FR / CC0 1.0
these weights Apache-2.0, with the attribution above

Limitations

  • Trained on Tatoeba, which is short everyday sentences. Expect it to be weakest on long, technical or literary text.
  • Neural MT translates each sentence independently, so a term may be rendered differently across a corpus. Check consistency if that matters to you.
  • Output should be reviewed by a speaker before being taught to learners.
Downloads last month
-
Safetensors
Model size
77M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for linguardia/opus-mt-en-jbo

Finetuned
(24)
this model

Dataset used to train linguardia/opus-mt-en-jbo