Instructions to use linguardia/opus-mt-en-kab with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use linguardia/opus-mt-en-kab with Transformers:
# Use a pipeline as a high-level helper # Warning: Pipeline type "translation" is no longer supported in transformers v5. # You must load the model directly (see below) or downgrade to v4.x with: # 'pip install "transformers<5.0.0' from transformers import pipeline pipe = pipeline("translation", model="linguardia/opus-mt-en-kab")# Load model directly from transformers import AutoTokenizer, AutoModelForSeq2SeqLM tokenizer = AutoTokenizer.from_pretrained("linguardia/opus-mt-en-kab") model = AutoModelForSeq2SeqLM.from_pretrained("linguardia/opus-mt-en-kab", device_map="auto") - Notebooks
- Google Colab
- Kaggle
English → Kabyle (opus-mt finetune)
Translates English into Kabyle. Finetuned by
Linguardia from Helsinki-NLP/opus-mt-en-mul on the
English↔Kabyle sentence pairs in Tatoeba.
Built to generate course material for a language that no major machine translator supports.
Usage
from transformers import AutoTokenizer, AutoModelForSeq2SeqLM
tok = AutoTokenizer.from_pretrained("linguardia/opus-mt-en-kab")
model = AutoModelForSeq2SeqLM.from_pretrained("linguardia/opus-mt-en-kab")
batch = tok([">>kab<< I wake up early."], return_tensors="pt")
out = model.generate(**batch, num_beams=4, max_length=64)
print(tok.batch_decode(out, skip_special_tokens=True))
The target language is selected by the >>kab<< prefix on the source
text. It was added as a new token rather than replacing one of the base
model's existing languages, so every language the base already handled still
works.
Training
| base | Helsinki-NLP/opus-mt-en-mul (77M params) |
| training pairs | 138,740 |
| held out | 1,000 dev, 137 gold |
| epochs | 2.0 |
| batch size | 32 |
| optimiser | Adafactor, lr 5e-5, 500 warmup steps |
| max sequence | 48 tokens (99.9% of source, 99.6% of target) |
| hardware | Apple M4, MPS |
Evaluation
Held-out Tatoeba (1000 pairs never seen in training):
| metric | score |
|---|---|
| BLEU | 26.2 |
| chrF | 50.9 |
Against 137 human translations of sentences that appear verbatim in the target corpus:
| metric | score |
|---|---|
| BLEU | 22.9 |
| chrF | 44.2 |
Translating the same 2,000-sentence course corpus:
| corpus | rows | distinct | vocab | words/sentence |
|---|---|---|---|---|
| this model | 2000 | 1999 (100.0%) | 2207 | 6.9 |
Training data and attribution
Sentence pairs come from Tatoeba, released under CC-BY 2.0 FR, with some sentences under CC0 1.0. Tatoeba's sentences are written by volunteers, and this model is a derivative of their work: if you use it, credit Tatoeba and its contributors.
Pairs were deduplicated on exact (English, Kabyle) text. No machine translation was used to create the training data at any point.
Licence chain
| layer | licence |
|---|---|
base model Helsinki-NLP/opus-mt-en-mul |
Apache-2.0 |
| training data (Tatoeba) | CC-BY 2.0 FR / CC0 1.0 |
| these weights | Apache-2.0, with the attribution above |
Limitations
- Trained on Tatoeba, which is short everyday sentences. Expect it to be weakest on long, technical or literary text.
- Neural MT translates each sentence independently, so a term may be rendered differently across a corpus. Check consistency if that matters to you.
- Output should be reviewed by a speaker before being taught to learners.
- Downloads last month
- -
Model tree for linguardia/opus-mt-en-kab
Base model
Helsinki-NLP/opus-mt-en-mul