Instructions to use ALJIACHI/Mizan-Rerank-v3 with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- sentence-transformers
How to use ALJIACHI/Mizan-Rerank-v3 with sentence-transformers:
from sentence_transformers import CrossEncoder model = CrossEncoder("ALJIACHI/Mizan-Rerank-v3", trust_remote_code=True) query = "Which planet is known as the Red Planet?" passages = [ "Venus is often called Earth's twin because of its similar size and proximity.", "Mars, known for its reddish appearance, is often referred to as the Red Planet.", "Jupiter, the largest planet in our solar system, has a prominent red spot.", "Saturn, famous for its rings, is sometimes mistaken for the Red Planet." ] scores = model.predict([(query, passage) for passage in passages]) print(scores) - Notebooks
- Google Colab
- Kaggle
Mizan-Rerank-v3
A compact, fast Arabic reranker that tells the passage that answers the question apart from passages that only look like they do.
Mizan-Rerank-v3 is a 306M-parameter Arabic cross-encoder. It was fine-tuned from Alibaba-NLP/gte-multilingual-reranker-base on 71k Arabic listwise groups. Each group pairs the correct passage with adversarial "trap" passages that share most of its words but change its meaning: a negated ruling, a swapped entity, a shifted number or date, an exception applied to the wrong case.
The result is a small model with the accuracy of a much larger one. It has about half the parameters of bge-reranker-v2-m3 (568M) and scores twice as many query–passage pairs per second, yet it ranks higher on average across our held-out Arabic benchmarks. It is built for Arabic search and RAG pipelines, especially over long, dense religious and encyclopedic text, where reranking latency matters.
Highlights
- Accurate for its size. Held-out mean nDCG@10 is 0.845 at 306M parameters. The 568M bge-reranker-v2-m3 scores 0.802, the gte base 0.783 and Mizan-Rerank-V2 0.723.
- Fast. On one RTX 4090 in fp16 it scores 1,228 pairs/s, 2.0× bge-reranker-v2-m3's 600 pairs/s, with 23% less peak GPU memory (see Efficiency).
- Picks the right passage first. Hit@1 on held-out long-context adversarial queries is 0.570, vs 0.336 for bge-reranker-v2-m3, 0.314 for gte base and 0.227 for Mizan-Rerank-V2.
- Rarely puts a contradiction on top. Across 2,666 held-out adversarial queries, the top-ranked passage is a contradicting trap 32.5% of the time. For bge-reranker-v2-m3 the figure is 56.8%; for Mizan-Rerank-V2, 73.0%.
- A drop-in upgrade over Mizan-Rerank-V2. It has the same architecture, size and speed, and scores higher on every benchmark.
Usage
Sentence Transformers
pip install -U sentence-transformers
from sentence_transformers import CrossEncoder
model = CrossEncoder("ALJIACHI/Mizan-Rerank-v3", trust_remote_code=True, max_length=3072)
query = "ما هي فوائد فيتامين د؟"
passages = [
"يساعد فيتامين د على امتصاص الكالسيوم وتقوية العظام، كما يدعم عمل الجهاز المناعي.",
"يستخدم فيتامين د في بعض الصناعات الغذائية كمادة مضافة لتدعيم الحليب.",
"أطلقت وزارة الزراعة حملة وطنية لزيادة الوعي بأهمية الزراعة العضوية.",
]
scores = model.predict([(query, passage) for passage in passages])
print(scores) # ≈ [0.875, 0.023, 0.010] (sigmoid relevance scores)
for hit in model.rank(query, passages, return_documents=True):
print(f"{hit['score']:.3f} {hit['text']}")
Transformers
import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer
tokenizer = AutoTokenizer.from_pretrained("ALJIACHI/Mizan-Rerank-v3")
model = AutoModelForSequenceClassification.from_pretrained(
"ALJIACHI/Mizan-Rerank-v3", trust_remote_code=True, torch_dtype=torch.float16
).to("cuda").eval()
query = "ما هي فوائد فيتامين د؟"
passages = [
"يساعد فيتامين د على امتصاص الكالسيوم وتقوية العظام، كما يدعم عمل الجهاز المناعي.",
"أطلقت وزارة الزراعة حملة وطنية لزيادة الوعي بأهمية الزراعة العضوية.",
]
with torch.inference_mode():
features = tokenizer([query] * len(passages), passages, padding=True, truncation=True,
max_length=3072, return_tensors="pt").to("cuda")
scores = torch.sigmoid(model(**features).logits.squeeze(-1).float())
for passage, score in sorted(zip(passages, scores.tolist()), key=lambda item: -item[1]):
print(f"{score:.3f} {passage}")
Practical notes
- Sequence length. The model was trained on query + passage inputs of up to 3,072 tokens and is benchmarked
here at 2,048. The architecture accepts 8,192 positions, but quality beyond 3,072 tokens has not been measured.
If your query + passage inputs fit in 512 tokens,
max_length=512is much faster. - Scores. The sigmoid of the single output logit is a relevance score. It works for ranking and for coarse cut-offs, but it is not a calibrated probability: tune any threshold on your own data.
trust_remote_code=Trueis required. The architecture code (modeling.py,configuration.py) comes from the gte base model and ships in this repository.
Evaluation
All four models were scored by the same script (benchmark/benchmark_rerankers.py)
with the same inputs: float32, max_length=2048, and the full candidate list of every query. Metrics use graded
gains (correct passage 3, partially correct passage 1, trap or irrelevant passage 0).
| Metric | Meaning |
|---|---|
| nDCG@10 | quality of the whole ranking |
| MRR@10 | 1 / rank of the correct passage |
| Hit@1 | share of queries where the correct passage is ranked first |
Held-out benchmarks
| Benchmark | Queries | Mizan-Rerank-v3 | Mizan-Rerank-V2 | gte-multilingual-reranker-base | bge-reranker-v2-m3 |
|---|---|---|---|---|---|
| MTEB NamaaMrTydi, unseen subset¹ | 573 | 0.8784 | 0.8220 | 0.8771 | 0.8897 |
| Arabic Hard Negatives, unseen subset¹ | 314 | 0.9200 | 0.8780 | 0.9062 | 0.9035 |
| Adversarial short (test)² | 602 | 0.8432 | 0.6965 | 0.7565 | 0.7938 |
| Adversarial long-context (test)² | 1,703 | 0.8280 | 0.6203 | 0.6747 | 0.6873 |
| Multi-LLM benchmark² | 361 | 0.7536 | 0.5998 | 0.6992 | 0.7365 |
| Average nDCG@10 | 0.8446 | 0.7233 | 0.7827 | 0.8022 | |
| Average MRR@10 | 0.7908 | 0.6312 | 0.7160 | 0.7427 | |
| Average Hit@1 | 0.6461 | 0.4221 | 0.5307 | 0.5764 |
¹ Queries whose question and correct passage never occur in the v3 training pool (hashes in
benchmark/seen_in_training.json).
² Internal test sets, not released. They share no queries or passages with the training data (verified by exact
match), and the long-context split is by source article.
Where the gain comes from: ranking the correct passage first. Hit@1 is where v3 differs most from the other models, especially when traps share most of their text with the answer:
Contradiction ranked first
For RAG the worst failure is not a missing answer but a passage that says the opposite ranked at the top: a negated ruling, a reversed condition, a wrong number. Every adversarial test query has one correct passage and several contradicting traps. The table shows how often a trap was ranked #1 (lower is better):
| Test set | Queries | Mizan-Rerank-v3 | Mizan-Rerank-V2 | gte-multilingual-reranker-base | bge-reranker-v2-m3 |
|---|---|---|---|---|---|
| Adversarial short | 602 | 20.1% (121) | 62.1% (374) | 51.3% (309) | 43.5% (262) |
| Adversarial long-context | 1,703 | 33.4% (568) | 76.4% (1,301) | 68.2% (1,161) | 62.0% (1,056) |
| Multi-LLM | 361 | 49.3% (178) | 75.3% (272) | 64.0% (231) | 54.0% (195) |
Across all 2,666 queries, v3 ranks a trap first 32.5% of the time. For bge-reranker-v2-m3 the figure is 56.8%; for gte base, 63.8%; for Mizan-Rerank-V2, 73.0%:
Reproduce with python compare_generated_test_models.py --dataset-file <test.jsonl> ... --chart contradiction_first.png
(in the training repository; the test sets are internal).
Full per-benchmark tables (nDCG@10 / MRR@10 / Hit@1), including full-set public scores
nDCG@10
| Model | NamaaMrTydi (full, 918) | NamaaMrTydi (unseen, 573) | Arabic Hard Neg. (full, 12,373)³ | Arabic Hard Neg. (unseen, 314) | Adversarial short | Adversarial long-ctx | Multi-LLM |
|---|---|---|---|---|---|---|---|
| Mizan-Rerank-v3 | 0.8885 | 0.8784 | 0.9103 | 0.9200 | 0.8432 | 0.8280 | 0.7536 |
| Mizan-Rerank-V2 | 0.8312 | 0.8220 | 0.8447 | 0.8780 | 0.6965 | 0.6203 | 0.5998 |
| gte-multilingual-reranker-base | 0.8874 | 0.8771 | 0.8960 | 0.9062 | 0.7565 | 0.6747 | 0.6992 |
| bge-reranker-v2-m3 | 0.8995 | 0.8897 | 0.9082 | 0.9035 | 0.7938 | 0.6873 | 0.7365 |
MRR@10
| Model | NamaaMrTydi (full) | NamaaMrTydi (unseen) | Arabic Hard Neg. (full)³ | Arabic Hard Neg. (unseen) | Adversarial short | Adversarial long-ctx | Multi-LLM |
|---|---|---|---|---|---|---|---|
| Mizan-Rerank-v3 | 0.8518 | 0.8385 | 0.8796 | 0.8926 | 0.7731 | 0.7658 | 0.6841 |
| Mizan-Rerank-V2 | 0.7755 | 0.7635 | 0.7919 | 0.8362 | 0.5952 | 0.4983 | 0.4629 |
| gte-multilingual-reranker-base | 0.8498 | 0.8362 | 0.8607 | 0.8743 | 0.6826 | 0.5765 | 0.6105 |
| bge-reranker-v2-m3 | 0.8662 | 0.8531 | 0.8771 | 0.8711 | 0.7374 | 0.5879 | 0.6643 |
Hit@1
| Model | NamaaMrTydi (full) | NamaaMrTydi (unseen) | Arabic Hard Neg. (full)³ | Arabic Hard Neg. (unseen) | Adversarial short | Adversarial long-ctx | Multi-LLM |
|---|---|---|---|---|---|---|---|
| Mizan-Rerank-v3 | 0.7691 | 0.7522 | 0.7882 | 0.8121 | 0.6146 | 0.5696 | 0.4820 |
| Mizan-Rerank-V2 | 0.6492 | 0.6335 | 0.6433 | 0.7134 | 0.3405 | 0.2267 | 0.1967 |
| gte-multilingual-reranker-base | 0.7549 | 0.7347 | 0.7600 | 0.7834 | 0.4618 | 0.3136 | 0.3601 |
| bge-reranker-v2-m3 | 0.7865 | 0.7627 | 0.7916 | 0.7866 | 0.5449 | 0.3365 | 0.4515 |
³ 97% of these queries occur in the v3 training pool, so the full-set score is not a fair measure for v3 and is not bolded or used in any average.
Efficiency
Every model scored the same 2,048 Arabic query–passage pairs from MTEB NamaaMrTydi. The run used one NVIDIA RTX
4090, float16, batch size 32 and max_length=512, after a warm-up
(benchmark/speed_benchmark.py):
| Model | Parameters | Throughput (pairs/s) | Peak GPU memory | Held-out mean nDCG@10 |
|---|---|---|---|---|
| Mizan-Rerank-v3 | 306M | 1,228 | 1.10 GiB | 0.845 |
| Mizan-Rerank-V2 | 306M | 1,219 | 1.10 GiB | 0.723 |
| gte-multilingual-reranker-base | 306M | 1,233 | 1.10 GiB | 0.783 |
| bge-reranker-v2-m3 | 568M | 600 | 1.43 GiB | 0.802 |
At the same latency budget, v3 can rerank twice as many candidates as bge-reranker-v2-m3. It reaches the highest held-out accuracy of the four models while running at the speed of the smallest.
Training
Data (71,044 listwise groups)
Every group is one query with one correct passage, up to 5 hard negatives and, for adversarial groups, one partially correct passage.
| Source | Groups | Share | What it teaches |
|---|---|---|---|
| Long-context adversarial (OpenWiki articles, LLM-generated) | 36,280 | 51% | find the answer inside long, dense passages, whether it sits at the beginning, middle or end |
| Short adversarial (LLM-generated) | 12,030 | 17% | reject near-duplicate traps in short passages |
| Arabic Mr. TyDi retrieval replay | 8,525 | 12% | keep general open-domain retrieval ability |
| Hadith question–passage pairs | 4,736 | 7% | classical religious text |
| Arabic news similarity pairs with LLM hard negatives | 4,737 | 7% | topical near-misses |
| Fatwa Q&A pairs with TF-IDF-mined negatives | 4,736 | 7% | jurisprudence rulings on nearby cases |
The adversarial negatives cover 12 trap types, each a minimal edit that changes the meaning while keeping most
of the words: negation_flip, exception_scope, entity_role_reversal, entity_swap, ordered_precedence,
temporal_boundary, numeric_boundary, quantifier_change, causal_direction, condition_swap,
conclusion_flip, attribution_shift.
Leakage controls: the adversarial train, validation and test splits are split by source article. Extra training data sharing a source article or a query with any validation or test set was removed. The three legacy sources were additionally filtered against every evaluation set by 8-word shingles. All held-out sets above were re-checked against the final training pool by exact query and passage match.
Objective
For each group the model scores all candidates jointly:
- Pointwise: binary cross-entropy, targets 1 / 0.5 / 0 for correct / partial / negative,
pos_weight=1.24, weight 0.5. - Pairwise: softplus margin loss
softplus(s_neg − s_pos + 0.2)for every negative, weight 1.0. - Partial order: the same margin loss for correct > partial > negative, weight 0.5 within the pairwise term.
Configuration
| Parameter | Value |
|---|---|
| Initialization | Alibaba-NLP/gte-multilingual-reranker-base |
| Max sequence length (query + passage) | 3,072 tokens |
| Candidates per group | 1 correct + 1 partial (if present) + up to 5 negatives |
| Batch | 4 groups/GPU × 8 accumulation × 2 GPUs = 64 groups per update |
| Optimizer | AdamW, lr 1e-6, weight decay 0.01, grad-norm clip 1.0 |
| Schedule | cosine, 10% warmup, 7 epochs (7,777 updates) |
| Precision | bf16, gradient checkpointing, SDPA attention |
| Hardware | 2 × NVIDIA RTX 4090 |
| Checkpoint selection | epoch with the best mean validation nDCG@10 (long-context + short adversarial) → epoch 6 |
Framework versions
Python 3.10.14 · PyTorch 2.8.0+cu126 · Transformers 4.55.4 · Sentence Transformers 5.4.1 · Accelerate 1.10.0 · Tokenizers 0.21.0
Citation
@software{Mizan_Rerank_v3_2026,
author = {Ali Aljiachi},
title = {Mizan-Rerank-v3: Adversarially Trained Arabic Long-Context Reranker},
year = {2026},
publisher = {Hugging Face},
url = {https://huggingface.co/ALJIACHI/Mizan-Rerank-v3}
}
License
Apache 2.0, the same license as the base model Alibaba-NLP/gte-multilingual-reranker-base.
- Downloads last month
- -
Model tree for ALJIACHI/Mizan-Rerank-v3
Base model
Alibaba-NLP/gte-multilingual-reranker-base





