Instructions to use pushkarsharma/wmt26-arabic-asian-mt with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- PEFT
How to use pushkarsharma/wmt26-arabic-asian-mt with PEFT:
Task type is invalid.
- Notebooks
- Google Colab
- Kaggle
WMT 2026 Low-Resource Arabic–Asian MT — NLP-IIT-Patna
Fine-tuned weights for our submission to the WMT 2026 shared task on Low-Resource Arabic–Asian Machine Translation: six directions (Arabic↔English, Arabic↔Hindi, Arabic↔Urdu) across three multilingual models, each direction trained separately.
Code, data preparation and evaluation: github.com/poskar-a/Arabic-translation-challenges
| Folder | Base model | Adaptation | Submitted as |
|---|---|---|---|
madlad/ |
google/madlad400-10b-mt |
LoRA, frozen 8-bit base | primary |
nllb/ |
facebook/nllb-200-3.3B |
full fine-tuning, bf16 | contrastive-1 |
gemmax2/ |
ModelSpace/GemmaX2-28-9B-v0.1 |
QLoRA, 4-bit NF4 | contrastive-2 |
Each folder holds one subfolder per direction — ar-en, en-ar, ar-hi, hi-ar,
ar-ur, ur-ar — with the tokenizer and trainer state alongside the weights. NLLB
folders contain full model weights; MADLAD and GemmaX2 folders contain LoRA adapters
to be applied on top of the base checkpoints above.
Official ranks
Blind challenge test set. Primary and contrastive systems are ranked on separate leaderboards, so the columns are not directly comparable.
| Direction | MADLAD-400 (primary) | NLLB-200 (contr-1) | GemmaX2 (contr-2) |
|---|---|---|---|
| en→ar | 1 | 4 | 3 |
| hi→ar | 1 | 3 | 2 |
| ur→ar | 1 | 3 | 2 |
| ar→en | 2 | 4 | 2 |
| ar→hi | 1 | 1 | 2 |
| ar→ur | 3 | 1 | 2 |
Usage
Download one direction rather than the whole repo (NLLB alone is 40 GB):
from huggingface_hub import snapshot_download
path = snapshot_download(
"pushkarsharma/wmt26-arabic-asian-mt",
allow_patterns="madlad/ar-en/*",
)
MADLAD-400 (primary). Prefix the source with the target-language tag:
from peft import PeftModel
from transformers import AutoModelForSeq2SeqLM, AutoTokenizer
base = AutoModelForSeq2SeqLM.from_pretrained("google/madlad400-10b-mt", device_map="auto")
model = PeftModel.from_pretrained(base, f"{path}/madlad/ar-en")
tok = AutoTokenizer.from_pretrained(f"{path}/madlad/ar-en")
inputs = tok("<2en> مرحبا بالعالم", return_tensors="pt").to(model.device)
print(tok.decode(model.generate(**inputs, num_beams=4)[0], skip_special_tokens=True))
NLLB-200. Set the source language and force the target BOS token
(arb_Arab, eng_Latn, hin_Deva, urd_Arab):
tok = AutoTokenizer.from_pretrained(f"{path}/nllb/ar-en")
model = AutoModelForSeq2SeqLM.from_pretrained(f"{path}/nllb/ar-en", device_map="auto")
tok.src_lang = "arb_Arab"
inputs = tok("مرحبا بالعالم", return_tensors="pt").to(model.device)
ids = model.generate(**inputs, forced_bos_token_id=tok.convert_tokens_to_ids("eng_Latn"), num_beams=4)
print(tok.decode(ids[0], skip_special_tokens=True))
GemmaX2-28-9B. Instruction prompt, adapter on a 4-bit base:
Translate the following text from Arabic to English:
{source}
Translation:
Decoding for all reported results: beam 4, no_repeat_ngram_size=3, repetition
penalty 1.3, length penalty 0.8, max_new_tokens = min(512, 2.5 × source length),
outputs normalised to Unicode NFC.
Training
One run per direction on a single RTX 6000 Ada (49 GB), seed 42, effective batch size 32, linear schedule with warmup, early stopping at patience 2.
| NLLB-200 | MADLAD-400 | GemmaX2-28-9B | |
|---|---|---|---|
| Adaptation | full FT | LoRA | QLoRA |
| Precision | bf16 | 8-bit base | 4-bit NF4, bf16 compute |
| Epochs | 5 | 5 | 3 |
| Learning rate | 5e-5 | 5e-5 | 2e-4 |
| LoRA r / α / dropout | — | 16 / 32 / 0.05 | 16 / 32 / 0.05 |
| Checkpoint selection | COMET-22 | eval loss | eval loss |
No heuristic normalisation was applied before tokenisation: the usual Arabic rules collapse Urdu characters that are contrastive, and each model's subword vocabulary was learned over unnormalised text.
License
Released under CC BY-NC 4.0, the most restrictive of the upstream licenses (NLLB-200 is CC BY-NC 4.0; MADLAD-400 is Apache 2.0; GemmaX2 inherits the Gemma license). Base-model terms apply to their respective derivatives.
- Downloads last month
- -
Model tree for pushkarsharma/wmt26-arabic-asian-mt
Base model
google/gemma-2-9b