AraGenre 8-Model Ensemble — Dev-Tuned Weights

A fixed-weight combination of 8 component sentence-encoder models' dev scores for hierarchical Arabic genre classification. This is not a single fine-tuned checkpoint — the repo contains a weights JSON and a combination script, not trained weights of its own.

Authors: Hassan Barmandah (NAMAA Community; Umm Al-Qura University), Israa Elhosiny (NAMAA Community), Yousra El-Ghawi (NAMAA Community), Omer Nacar (NAMAA Community)

⚠️ Generalization Note

This ensemble's development-set score (0.9569 hierarchical F1) is not representative of real-world performance. Per the project's system-description paper, this entire lineage of fine-tuned/ensembled sentence encoders — which scored well on the 110-item, 6-genre AraGenre dev set — collapsed to 0.22–0.44 hierarchical F1 on the actual 27,972-item hidden test set (74 specific genres under 6 broad genres). Its component weights were fit via a grid search directly against dev labels, so even the 0.9569 number reflects fitting to dev, not just evaluation on it.

The system that actually won for this team — 0.7013 hierarchical F1, 3rd of 18 teams on the official CodaBench leaderboard — was a separate, zero-shot DeepSeek-LLM pipeline with no fine-tuning at all (stage2_llm_zeroshot_pipeline/ in the project repo). This artifact is not that system. It is released here for transparency and reproducibility of the project's full experimental record, not as a recommended production classifier.

Approach

Combines 8 component models' softmax-normalized similarity scores using fixed, dev-grid-searched weights:

bge-m3-zeroshot: 0.053              bge-m3-augmented-defs: 0.158
e5-large-cosine-8ep: 0.0             e5-large-cosine-10ep: 0.263
e5-large-mnrl-xgenre-lite: 0.211     e5-large-mnrl-augmented-defs: 0.105
e5-large-mnrl-xgenre-phase1: 0.158   e5-large-multiseed-ensemble: 0.053

Each component's raw similarity scores are softmax-normalized first (so cross-model score scales are comparable), then combined with the weights above.

Training data

None directly — this is a weight recipe over pre-scored component models. The component weights were selected via grid search against dev_gold.json labels.

Usage

Requires cached dev score files from running all 8 component scripts first (stage1_encoder_finetuning/scores/*_dev_scores.json), then:

python ensemble_8model_devtuned.py

See the project repository for the full script and component-model requirements.

Citation

If you use this work, please cite our system-description paper:

@inproceedings{barmandah-etal-2026-namaa,
  title = {NAMAA at AraGenre 2026: From Encoder Baselines to Self-Consistent LLM Ensembling for Hierarchical Arabic Genre Classification},
  author = {Barmandah, Hassan and Elhosiny, Israa and El-Ghawi, Yousra and Nacar, Omer},
  booktitle = {Proceedings of the 4th Arabic Natural Language Processing Conference (ArabicNLP 2026)},
  address = {Budapest, Hungary},
  publisher = {Association for Computational Linguistics},
  year = {2026},
}

Please also cite the AraGenre 2026 shared task overview paper:

@inproceedings{elhaj-etal-2026-aragenre,
  title = {AraGenre 2026: A Hierarchical Definition-Guided Arabic Genre Classification Shared Task},
  author = {El-Haj, Mo and Ezzini, Saad and Abudalfa, Shadi and Lamsiyah, Salima and Jarrar, Mustafa},
  booktitle = {Proceedings of the 4th Arabic Natural Language Processing Conference (ArabicNLP 2026)},
  address = {Budapest, Hungary},
  publisher = {Association for Computational Linguistics},
  year = {2026},
}

License

Apache 2.0

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Collection including HassanB4/aragenre-8model-ensemble-devtuned

Evaluation results

  • Hierarchical Macro F1 (DEVELOPMENT SET, not a test-set metric) on AraGenre 2026 Development Set
    self-reported
    0.957