French–Serer OPUS-MT Full Fine-Tuning

Part of a benchmark of six configurations for French→Serer neural machine translation. Serer is a critically low-resource Niger-Congo language (~1.2M speakers, Senegal/Gambia), phylogenetically close to the well-resourced Wolof.

Model summary

  • Experiment ID: E
  • Kind: final_translation_model
  • Direction: French → Serer
  • Base model: Helsinki-NLP/opus-mt-fr-en (revision: main)
  • Best checkpoint: french_serer_opusmt_full_ft-epoch=09-val_bleu=15.6038.ckpt
  • Random seed: 42

⚠️ Loading this model — do not use a plain from_pretrained() call

This model has asymmetric encoder/decoder vocabularies: the French encoder keeps the original 59,514-token opus-mt-fr-en vocabulary, while the Serer decoder was resized to 8,000 tokens. MarianModel builds both from a single config.vocab_size at construction time, so a plain MarianMTModel.from_pretrained("Fallovski/french-serer-opusmt-full-ft") call will either crash (CUDA error: device-side assert triggered, out-of-bounds embedding lookup) or silently return incoherent output if you pass ignore_mismatched_sizes=True without the extra step below, because the encoder embeddings end up randomly initialized instead of loaded from the checkpoint.

The correct French encoder weights are saved in this repository, under their own key (model.encoder.embed_tokens.weight) — they just need to be re-attached manually after loading, since MarianModel's architecture cannot represent this asymmetry on its own. This is a known limitation of the export, not data loss: every weight needed is present and correct.

from transformers import MarianMTModel, MarianTokenizer
from safetensors.torch import load_file
from huggingface_hub import hf_hub_download
import torch.nn as nn

repo = "Fallovski/french-serer-opusmt-full-ft"
model = MarianMTModel.from_pretrained(repo, ignore_mismatched_sizes=True).eval()
tokenizer = MarianTokenizer.from_pretrained(repo)

# Required: MarianModel builds encoder/decoder from one shared config.vocab_size
# at construction time, so it cannot natively represent this model's asymmetric
# vocabularies (French encoder: 59,514 tokens; Serer decoder: 8,000 tokens).
# Re-attach the correct French encoder embeddings, saved under their own key:
weights_path = hf_hub_download(repo, "model.safetensors")
encoder_weight = load_file(weights_path)["model.encoder.embed_tokens.weight"]
model.model.encoder.embed_tokens = nn.Embedding.from_pretrained(encoder_weight, padding_idx=0)

# Decode Serer output with the target sentencepiece model (not `tokenizer`,
# which only covers the French source side):
import sentencepiece as spm
target_sp = spm.SentencePieceProcessor(model_file=hf_hub_download(repo, "serer_spm.model"))

Verified working end-to-end (2026-09-13): generate() produces valid output ids entirely within the 8,000-token Serer vocabulary, no crash, no ignore_mismatched_sizes warnings left unresolved.

Hyperparameters

  • LR: 2e-05
  • NUM_EPOCHS: 15
  • TARGET_VOCAB_SIZE: 8000
  • WARMUP_STEPS: 500

This model has no Atlantic-family prior (its pretraining does not include Wolof or any close relative of Serer), unlike the NLLB-based configurations in this project. In our evaluation, this configuration was judged the most linguistically reliable by a native Serer-speaking expert despite a lower BLEU score than NLLB-based alternatives — see the paper for the full BLEU/quality decorrelation analysis.

Intended use

Research on French-to-Serer machine translation on a corpus that is ~90% religious (Bible) register, ~10% educational glossaries, primarily Siin dialect. Not validated for legal, medical, emergency, or fully autonomous publication use. Private repository — not intended for public deployment in its current state.

Evaluation

Metric Value
BLEU (test, beam=5) 17.7427
chrF n/a
ROUGE-1 0.4086
ROUGE-L 0.359
BERTScore-F1 0.8748
Test loss 3.379

Evaluated on the held-out test split (2890 sentence pairs, SHA-256 of the split: 01d14d982a3c0cce172b5099e2db064d05bdef89677fe72a45f56715d6364ee2). Metrics were computed with the project's own evaluation scripts (not copied from the manuscript without independent reproduction); the training and evaluation code is kept in a private repository, available on request.

Training data and rights

Parallel corpus of 23113 train / 2889 val / 2890 test French–Serer sentence pairs, built primarily from religious texts (Bible, 90%) and educational glossaries (10%), predominantly Siin dialect. Preprocessing: Unicode normalization, exact-duplicate removal, length-ratio filtering (1:3–3:1). Document-level splitting was not possible (no document identifiers available); the split is at the sentence level with a fixed seed. Full provenance, licensing, and consent documentation are kept in a private dataset card, available on request, prior to any public release.

Downloads last month
56
Safetensors
Model size
82.8M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support