AIN β€” English ↔ Igbo Translation Model

This model is a fine-tuned version of facebook/nllb-200-distilled-600M, adapted for English -- Igbo translation as part of the aAIN project β€” a speech-to-speech translation system for English and Nigerian languages (Yoruba, Hausa, Igbo).

Model description

The base NLLB-200-distilled-600M model was fine-tuned on 27,097 Igbo-English sentence pairs extracted and cleaned from a bilingual dictionary dataset. Fine-tuning was aimed at improving translation quality specifically for Igbo, a low-resource language with limited high-quality parallel data available for training.

Training results

Metric Before fine-tuning After fine-tuning
BLEU score 14.06 26.86

BLEU improved by +12.8 points on held-out test data β€” nearly double the baseline score.

Intended use

This model is intended for English ↔ Igbo text translation, as part of a larger speech-to-speech pipeline (ASR β†’ translation β†’ TTS). It can be used standalone for text translation via the transformers library.

Limitations

  • BLEU 26.86 indicates decent but imperfect translation quality. Expect generally coherent output, but not publication-grade accuracy on every sentence.
  • Some common conversational phrases (e.g. certain greetings) may occasionally be returned untranslated or copied verbatim from the source. This is a known limitation observed during testing and an area for improvement in future versions.
  • The training data was sourced from a dictionary dataset, which may bias the model toward more formal or literal phrasing rather than casual/conversational Igbo.
  • This model does not include Igbo text-to-speech; a separate TTS component is planned for a future release.

How to use

from transformers import AutoModelForSeq2SeqLM, AutoTokenizer

model_name = "Ntui-Chinua/aAIN"
tokenizer = AutoTokenizer.from_pretrained(model_name)
model = AutoModelForSeq2SeqLM.from_pretrained(model_name)

tokenizer.src_lang = "eng_Latn"
inputs = tokenizer("I am going to the market.", return_tensors="pt")
forced_bos_token_id = tokenizer.convert_tokens_to_ids("ibo_Latn")

generated_tokens = model.generate(**inputs, forced_bos_token_id=forced_bos_token_id, max_length=128)
print(tokenizer.batch_decode(generated_tokens, skip_special_tokens=True)[0])

Training data

27,097 English-Igbo sentence pairs extracted from a bilingual dictionary dataset.

License

Apache 2.0

Downloads last month
14
Safetensors
Model size
0.6B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support

Model tree for Ntui-Chinua/aAIN

Finetuned
(348)
this model