French–Serer OPUS-MT Full Fine-Tuning
Part of a benchmark of six configurations for French→Serer neural machine translation. Serer is a critically low-resource Niger-Congo language (~1.2M speakers, Senegal/Gambia), phylogenetically close to the well-resourced Wolof.
Model summary
- Experiment ID: E
- Kind: final_translation_model
- Direction: French → Serer
- Base model:
Helsinki-NLP/opus-mt-fr-en(revision:main) - Best checkpoint:
french_serer_opusmt_full_ft-epoch=09-val_bleu=15.6038.ckpt - Random seed: 42
⚠️ Loading this model — do not use a plain from_pretrained() call
This model has asymmetric encoder/decoder vocabularies: the French encoder keeps
the original 59,514-token opus-mt-fr-en vocabulary, while the Serer decoder was
resized to 8,000 tokens. MarianModel builds both from a single config.vocab_size
at construction time, so a plain MarianMTModel.from_pretrained("Fallovski/french-serer-opusmt-full-ft")
call will either crash (CUDA error: device-side assert triggered, out-of-bounds
embedding lookup) or silently return incoherent output if you pass
ignore_mismatched_sizes=True without the extra step below, because the encoder
embeddings end up randomly initialized instead of loaded from the checkpoint.
The correct French encoder weights are saved in this repository, under their own
key (model.encoder.embed_tokens.weight) — they just need to be re-attached manually
after loading, since MarianModel's architecture cannot represent this asymmetry on
its own. This is a known limitation of the export, not data loss: every weight needed
is present and correct.
from transformers import MarianMTModel, MarianTokenizer
from safetensors.torch import load_file
from huggingface_hub import hf_hub_download
import torch.nn as nn
repo = "Fallovski/french-serer-opusmt-full-ft"
model = MarianMTModel.from_pretrained(repo, ignore_mismatched_sizes=True).eval()
tokenizer = MarianTokenizer.from_pretrained(repo)
# Required: MarianModel builds encoder/decoder from one shared config.vocab_size
# at construction time, so it cannot natively represent this model's asymmetric
# vocabularies (French encoder: 59,514 tokens; Serer decoder: 8,000 tokens).
# Re-attach the correct French encoder embeddings, saved under their own key:
weights_path = hf_hub_download(repo, "model.safetensors")
encoder_weight = load_file(weights_path)["model.encoder.embed_tokens.weight"]
model.model.encoder.embed_tokens = nn.Embedding.from_pretrained(encoder_weight, padding_idx=0)
# Decode Serer output with the target sentencepiece model (not `tokenizer`,
# which only covers the French source side):
import sentencepiece as spm
target_sp = spm.SentencePieceProcessor(model_file=hf_hub_download(repo, "serer_spm.model"))
Verified working end-to-end (2026-09-13): generate() produces valid output ids
entirely within the 8,000-token Serer vocabulary, no crash, no ignore_mismatched_sizes
warnings left unresolved.
Hyperparameters
LR: 2e-05NUM_EPOCHS: 15TARGET_VOCAB_SIZE: 8000WARMUP_STEPS: 500
This model has no Atlantic-family prior (its pretraining does not include Wolof or any close relative of Serer), unlike the NLLB-based configurations in this project. In our evaluation, this configuration was judged the most linguistically reliable by a native Serer-speaking expert despite a lower BLEU score than NLLB-based alternatives — see the paper for the full BLEU/quality decorrelation analysis.
Intended use
Research on French-to-Serer machine translation on a corpus that is ~90% religious (Bible) register, ~10% educational glossaries, primarily Siin dialect. Not validated for legal, medical, emergency, or fully autonomous publication use. Private repository — not intended for public deployment in its current state.
Evaluation
| Metric | Value |
|---|---|
| BLEU (test, beam=5) | 17.7427 |
| chrF | n/a |
| ROUGE-1 | 0.4086 |
| ROUGE-L | 0.359 |
| BERTScore-F1 | 0.8748 |
| Test loss | 3.379 |
Evaluated on the held-out test split (2890 sentence pairs, SHA-256 of the split:
01d14d982a3c0cce172b5099e2db064d05bdef89677fe72a45f56715d6364ee2). Metrics were computed with the project's own evaluation scripts
(not copied from the manuscript without independent reproduction); the training and
evaluation code is kept in a private repository, available on request.
Training data and rights
Parallel corpus of 23113 train / 2889 val / 2890 test
French–Serer sentence pairs, built primarily from religious texts (Bible, 90%) and
educational glossaries (10%), predominantly Siin dialect. Preprocessing: Unicode
normalization, exact-duplicate removal, length-ratio filtering (1:3–3:1). Document-level
splitting was not possible (no document identifiers available); the split is at the
sentence level with a fixed seed. Full provenance, licensing, and consent documentation
are kept in a private dataset card, available on request, prior to any public release.
- Downloads last month
- 56