NLLB-200 600M Spanish ↔ Nasa Yuwe

Neural machine translation between Spanish and Nasa Yuwe, a Colombian Indigenous language (Nasa people, Cauca). The model is bidirectional: it translates Spanish to Nasa Yuwe and Nasa Yuwe to Spanish.

Fine-tune of facebook/nllb-200-distilled-600M from the thesis Data-Centric Strategies for Low-Resource Neural Machine Translation in Colombian Indigenous Languages (Universidad de los Andes, FLAG Lab).

Languages

Role Language Tokenizer code ISO 639-3
Source / target Spanish spa_Latn spa
Source / target Nasa Yuwe nas_Latn pbb

Results

SacreBLEU and chrF2 on the held-out test partition.

Direction BLEU chrF2
Spanish to Nasa Yuwe 3.37 18.37
Nasa Yuwe to Spanish 2.56 17.69

Usage

from transformers import AutoModelForSeq2SeqLM, NllbTokenizer

tok = NllbTokenizer.from_pretrained("Flaglab/nllb-600M-nasayuwe-esp")
model = AutoModelForSeq2SeqLM.from_pretrained("Flaglab/nllb-600M-nasayuwe-esp")

# Spanish to Nasa Yuwe
tok.src_lang = "spa_Latn"
inputs = tok("Buenos días, ¿cómo estás?", return_tensors="pt")
out = model.generate(**inputs, forced_bos_token_id=tok.convert_tokens_to_ids("nas_Latn"))
print(tok.batch_decode(out, skip_special_tokens=True))

Training

  • Base model: facebook/nllb-200-distilled-600M
  • Strategy: Expanded bilingual corpus (full variant).
  • Tokenizer: NLLB SentencePiece extended with a custom Nasa Yuwe vocabulary.

License and intended use

Released under CC-BY-NC 4.0, inherited from NLLB-200: non-commercial use only. Intended for research on machine translation for Colombian Indigenous languages. Outputs have not been validated by human speakers.

Downloads last month
28
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Flaglab/nllb-600M-nasayuwe-esp

Finetuned
(329)
this model

Collection including Flaglab/nllb-600M-nasayuwe-esp