AraGenre Fusion Ensemble (Multi-Embedding Fusion v3.0)
A multi-embedding k-NN + rule ensemble fusing E5-large, multilingual-mpnet, and Arabic-BERT centroids for hierarchical Arabic genre classification. Not a single fine-tuned checkpoint — this repo is a centroids/config system, not a trained model weight file.
Authors: Hassan Barmandah (NAMAA Community; Umm Al-Qura University), Israa Elhosiny (NAMAA Community), Yousra El-Ghawi (NAMAA Community), Omer Nacar (NAMAA Community)
⚠️ Generalization Note
This ensemble's development-set score (0.9255 hierarchical F1) is not representative of real-world performance. Per the project's system-description paper, this entire lineage of fine-tuned/ensembled sentence encoders — which scored well on the 110-item, 6-genre AraGenre dev set — collapsed to 0.22–0.44 hierarchical F1 on the actual 27,972-item hidden test set (74 specific genres under 6 broad genres). Its fusion weights were grid-searched directly against dev_gold labels, and its genre-reconciliation logic was also patched using dev_gold.json, so even the 0.9255 number reflects fitting to dev, not just evaluation on it.
The system that actually won for this team — 0.7013 hierarchical F1, 3rd of 18 teams on the official CodaBench leaderboard — was a separate, zero-shot DeepSeek-LLM pipeline with no fine-tuning at all (stage2_llm_zeroshot_pipeline/ in the project repo). This artifact is not that system. It is released here for transparency and reproducibility of the project's full experimental record, not as a recommended production classifier.
Approach
Rule + multi-embedding k-NN ensemble: E5-large + multilingual-mpnet + Arabic-BERT centroids, linguistic regex features, grid-searched fusion weights, two-stage broad → specific prediction. Extracted from the project's fusion-ensemble notebook, representing the last stage in its lineage before later versions introduced pattern rules hardcoded to specific dev example IDs (dev-set memorization, deliberately excluded from the project's results table).
Base models
E5-large + multilingual-mpnet + Arabic-BERT, combined via embedding fusion — not a single fine-tuned model.
Training data
AraGenre TRAIN genres (7) plus multi-model embedding fusion; fusion weights were grid-searched against dev labels (see Generalization Note above).
Usage
python multiembedding_fusion_ensemble_v3.py
Expects train/dev/dev_genre_definitions.json plus the synthetic-data JSON files (available as the aragenre-synthetic-training-data Hugging Face dataset).
Citation
If you use this work, please cite our system-description paper:
@inproceedings{barmandah-etal-2026-namaa,
title = {NAMAA at AraGenre 2026: From Encoder Baselines to Self-Consistent LLM Ensembling for Hierarchical Arabic Genre Classification},
author = {Barmandah, Hassan and Elhosiny, Israa and El-Ghawi, Yousra and Nacar, Omer},
booktitle = {Proceedings of the 4th Arabic Natural Language Processing Conference (ArabicNLP 2026)},
address = {Budapest, Hungary},
publisher = {Association for Computational Linguistics},
year = {2026},
}
Please also cite the AraGenre 2026 shared task overview paper:
@inproceedings{elhaj-etal-2026-aragenre,
title = {AraGenre 2026: A Hierarchical Definition-Guided Arabic Genre Classification Shared Task},
author = {El-Haj, Mo and Ezzini, Saad and Abudalfa, Shadi and Lamsiyah, Salima and Jarrar, Mustafa},
booktitle = {Proceedings of the 4th Arabic Natural Language Processing Conference (ArabicNLP 2026)},
address = {Budapest, Hungary},
publisher = {Association for Computational Linguistics},
year = {2026},
}
License
Apache 2.0
Collection including HassanB4/aragenre-multiembedding-fusion-ensemble-v3
Evaluation results
- Hierarchical Macro F1 (DEVELOPMENT SET, not a test-set metric) on AraGenre 2026 Development Setself-reported0.925