Mizan-Rerank-v3

A compact, fast Arabic reranker that tells the passage that answers the question apart from passages that only look like they do.

Parameters Speed Language Context License

Mizan-Rerank-v3 is a 306M-parameter Arabic cross-encoder. It was fine-tuned from Alibaba-NLP/gte-multilingual-reranker-base on 71k Arabic listwise groups. Each group pairs the correct passage with adversarial "trap" passages that share most of its words but change its meaning: a negated ruling, a swapped entity, a shifted number or date, an exception applied to the wrong case.

The result is a small model with the accuracy of a much larger one. It has about half the parameters of bge-reranker-v2-m3 (568M) and scores twice as many query–passage pairs per second, yet it ranks higher on average across our held-out Arabic benchmarks. It is built for Arabic search and RAG pipelines, especially over long, dense religious and encyclopedic text, where reranking latency matters.

Summary

Highlights

  • Accurate for its size. Held-out mean nDCG@10 is 0.845 at 306M parameters. The 568M bge-reranker-v2-m3 scores 0.802, the gte base 0.783 and Mizan-Rerank-V2 0.723.
  • Fast. On one RTX 4090 in fp16 it scores 1,228 pairs/s, 2.0× bge-reranker-v2-m3's 600 pairs/s, with 23% less peak GPU memory (see Efficiency).
  • Picks the right passage first. Hit@1 on held-out long-context adversarial queries is 0.570, vs 0.336 for bge-reranker-v2-m3, 0.314 for gte base and 0.227 for Mizan-Rerank-V2.
  • Rarely puts a contradiction on top. Across 2,666 held-out adversarial queries, the top-ranked passage is a contradicting trap 32.5% of the time. For bge-reranker-v2-m3 the figure is 56.8%; for Mizan-Rerank-V2, 73.0%.
  • A drop-in upgrade over Mizan-Rerank-V2. It has the same architecture, size and speed, and scores higher on every benchmark.

Usage

Sentence Transformers

pip install -U sentence-transformers
from sentence_transformers import CrossEncoder

model = CrossEncoder("ALJIACHI/Mizan-Rerank-v3", trust_remote_code=True, max_length=3072)

query = "ما هي فوائد فيتامين د؟"
passages = [
    "يساعد فيتامين د على امتصاص الكالسيوم وتقوية العظام، كما يدعم عمل الجهاز المناعي.",
    "يستخدم فيتامين د في بعض الصناعات الغذائية كمادة مضافة لتدعيم الحليب.",
    "أطلقت وزارة الزراعة حملة وطنية لزيادة الوعي بأهمية الزراعة العضوية.",
]

scores = model.predict([(query, passage) for passage in passages])
print(scores)  # ≈ [0.875, 0.023, 0.010]  (sigmoid relevance scores)

for hit in model.rank(query, passages, return_documents=True):
    print(f"{hit['score']:.3f}  {hit['text']}")

Transformers

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("ALJIACHI/Mizan-Rerank-v3")
model = AutoModelForSequenceClassification.from_pretrained(
    "ALJIACHI/Mizan-Rerank-v3", trust_remote_code=True, torch_dtype=torch.float16
).to("cuda").eval()

query = "ما هي فوائد فيتامين د؟"
passages = [
    "يساعد فيتامين د على امتصاص الكالسيوم وتقوية العظام، كما يدعم عمل الجهاز المناعي.",
    "أطلقت وزارة الزراعة حملة وطنية لزيادة الوعي بأهمية الزراعة العضوية.",
]

with torch.inference_mode():
    features = tokenizer([query] * len(passages), passages, padding=True, truncation=True,
                         max_length=3072, return_tensors="pt").to("cuda")
    scores = torch.sigmoid(model(**features).logits.squeeze(-1).float())

for passage, score in sorted(zip(passages, scores.tolist()), key=lambda item: -item[1]):
    print(f"{score:.3f}  {passage}")

Practical notes

  • Sequence length. The model was trained on query + passage inputs of up to 3,072 tokens and is benchmarked here at 2,048. The architecture accepts 8,192 positions, but quality beyond 3,072 tokens has not been measured. If your query + passage inputs fit in 512 tokens, max_length=512 is much faster.
  • Scores. The sigmoid of the single output logit is a relevance score. It works for ranking and for coarse cut-offs, but it is not a calibrated probability: tune any threshold on your own data.
  • trust_remote_code=True is required. The architecture code (modeling.py, configuration.py) comes from the gte base model and ships in this repository.

Evaluation

All four models were scored by the same script (benchmark/benchmark_rerankers.py) with the same inputs: float32, max_length=2048, and the full candidate list of every query. Metrics use graded gains (correct passage 3, partially correct passage 1, trap or irrelevant passage 0).

Metric Meaning
nDCG@10 quality of the whole ranking
MRR@10 1 / rank of the correct passage
Hit@1 share of queries where the correct passage is ranked first

Held-out benchmarks

Benchmark Queries Mizan-Rerank-v3 Mizan-Rerank-V2 gte-multilingual-reranker-base bge-reranker-v2-m3
MTEB NamaaMrTydi, unseen subset¹ 573 0.8784 0.8220 0.8771 0.8897
Arabic Hard Negatives, unseen subset¹ 314 0.9200 0.8780 0.9062 0.9035
Adversarial short (test)² 602 0.8432 0.6965 0.7565 0.7938
Adversarial long-context (test)² 1,703 0.8280 0.6203 0.6747 0.6873
Multi-LLM benchmark² 361 0.7536 0.5998 0.6992 0.7365
Average nDCG@10 0.8446 0.7233 0.7827 0.8022
Average MRR@10 0.7908 0.6312 0.7160 0.7427
Average Hit@1 0.6461 0.4221 0.5307 0.5764

¹ Queries whose question and correct passage never occur in the v3 training pool (hashes in benchmark/seen_in_training.json). ² Internal test sets, not released. They share no queries or passages with the training data (verified by exact match), and the long-context split is by source article.

Public benchmarks

Adversarial benchmarks

Where the gain comes from: ranking the correct passage first. Hit@1 is where v3 differs most from the other models, especially when traps share most of their text with the answer:

Hit@1

Contradiction ranked first

For RAG the worst failure is not a missing answer but a passage that says the opposite ranked at the top: a negated ruling, a reversed condition, a wrong number. Every adversarial test query has one correct passage and several contradicting traps. The table shows how often a trap was ranked #1 (lower is better):

Test set Queries Mizan-Rerank-v3 Mizan-Rerank-V2 gte-multilingual-reranker-base bge-reranker-v2-m3
Adversarial short 602 20.1% (121) 62.1% (374) 51.3% (309) 43.5% (262)
Adversarial long-context 1,703 33.4% (568) 76.4% (1,301) 68.2% (1,161) 62.0% (1,056)
Multi-LLM 361 49.3% (178) 75.3% (272) 64.0% (231) 54.0% (195)

Across all 2,666 queries, v3 ranks a trap first 32.5% of the time. For bge-reranker-v2-m3 the figure is 56.8%; for gte base, 63.8%; for Mizan-Rerank-V2, 73.0%:

Contradiction ranked first

Reproduce with python compare_generated_test_models.py --dataset-file <test.jsonl> ... --chart contradiction_first.png (in the training repository; the test sets are internal).

Full per-benchmark tables (nDCG@10 / MRR@10 / Hit@1), including full-set public scores

nDCG@10

Model NamaaMrTydi (full, 918) NamaaMrTydi (unseen, 573) Arabic Hard Neg. (full, 12,373)³ Arabic Hard Neg. (unseen, 314) Adversarial short Adversarial long-ctx Multi-LLM
Mizan-Rerank-v3 0.8885 0.8784 0.9103 0.9200 0.8432 0.8280 0.7536
Mizan-Rerank-V2 0.8312 0.8220 0.8447 0.8780 0.6965 0.6203 0.5998
gte-multilingual-reranker-base 0.8874 0.8771 0.8960 0.9062 0.7565 0.6747 0.6992
bge-reranker-v2-m3 0.8995 0.8897 0.9082 0.9035 0.7938 0.6873 0.7365

MRR@10

Model NamaaMrTydi (full) NamaaMrTydi (unseen) Arabic Hard Neg. (full)³ Arabic Hard Neg. (unseen) Adversarial short Adversarial long-ctx Multi-LLM
Mizan-Rerank-v3 0.8518 0.8385 0.8796 0.8926 0.7731 0.7658 0.6841
Mizan-Rerank-V2 0.7755 0.7635 0.7919 0.8362 0.5952 0.4983 0.4629
gte-multilingual-reranker-base 0.8498 0.8362 0.8607 0.8743 0.6826 0.5765 0.6105
bge-reranker-v2-m3 0.8662 0.8531 0.8771 0.8711 0.7374 0.5879 0.6643

Hit@1

Model NamaaMrTydi (full) NamaaMrTydi (unseen) Arabic Hard Neg. (full)³ Arabic Hard Neg. (unseen) Adversarial short Adversarial long-ctx Multi-LLM
Mizan-Rerank-v3 0.7691 0.7522 0.7882 0.8121 0.6146 0.5696 0.4820
Mizan-Rerank-V2 0.6492 0.6335 0.6433 0.7134 0.3405 0.2267 0.1967
gte-multilingual-reranker-base 0.7549 0.7347 0.7600 0.7834 0.4618 0.3136 0.3601
bge-reranker-v2-m3 0.7865 0.7627 0.7916 0.7866 0.5449 0.3365 0.4515

³ 97% of these queries occur in the v3 training pool, so the full-set score is not a fair measure for v3 and is not bolded or used in any average.

Efficiency

Every model scored the same 2,048 Arabic query–passage pairs from MTEB NamaaMrTydi. The run used one NVIDIA RTX 4090, float16, batch size 32 and max_length=512, after a warm-up (benchmark/speed_benchmark.py):

Model Parameters Throughput (pairs/s) Peak GPU memory Held-out mean nDCG@10
Mizan-Rerank-v3 306M 1,228 1.10 GiB 0.845
Mizan-Rerank-V2 306M 1,219 1.10 GiB 0.723
gte-multilingual-reranker-base 306M 1,233 1.10 GiB 0.783
bge-reranker-v2-m3 568M 600 1.43 GiB 0.802

Accuracy vs. speed

At the same latency budget, v3 can rerank twice as many candidates as bge-reranker-v2-m3. It reaches the highest held-out accuracy of the four models while running at the speed of the smallest.

Training

Data (71,044 listwise groups)

Every group is one query with one correct passage, up to 5 hard negatives and, for adversarial groups, one partially correct passage.

Source Groups Share What it teaches
Long-context adversarial (OpenWiki articles, LLM-generated) 36,280 51% find the answer inside long, dense passages, whether it sits at the beginning, middle or end
Short adversarial (LLM-generated) 12,030 17% reject near-duplicate traps in short passages
Arabic Mr. TyDi retrieval replay 8,525 12% keep general open-domain retrieval ability
Hadith question–passage pairs 4,736 7% classical religious text
Arabic news similarity pairs with LLM hard negatives 4,737 7% topical near-misses
Fatwa Q&A pairs with TF-IDF-mined negatives 4,736 7% jurisprudence rulings on nearby cases

The adversarial negatives cover 12 trap types, each a minimal edit that changes the meaning while keeping most of the words: negation_flip, exception_scope, entity_role_reversal, entity_swap, ordered_precedence, temporal_boundary, numeric_boundary, quantifier_change, causal_direction, condition_swap, conclusion_flip, attribution_shift.

Leakage controls: the adversarial train, validation and test splits are split by source article. Extra training data sharing a source article or a query with any validation or test set was removed. The three legacy sources were additionally filtered against every evaluation set by 8-word shingles. All held-out sets above were re-checked against the final training pool by exact query and passage match.

Objective

For each group the model scores all candidates jointly:

  • Pointwise: binary cross-entropy, targets 1 / 0.5 / 0 for correct / partial / negative, pos_weight=1.24, weight 0.5.
  • Pairwise: softplus margin loss softplus(s_neg − s_pos + 0.2) for every negative, weight 1.0.
  • Partial order: the same margin loss for correct > partial > negative, weight 0.5 within the pairwise term.

Configuration

Parameter Value
Initialization Alibaba-NLP/gte-multilingual-reranker-base
Max sequence length (query + passage) 3,072 tokens
Candidates per group 1 correct + 1 partial (if present) + up to 5 negatives
Batch 4 groups/GPU × 8 accumulation × 2 GPUs = 64 groups per update
Optimizer AdamW, lr 1e-6, weight decay 0.01, grad-norm clip 1.0
Schedule cosine, 10% warmup, 7 epochs (7,777 updates)
Precision bf16, gradient checkpointing, SDPA attention
Hardware 2 × NVIDIA RTX 4090
Checkpoint selection epoch with the best mean validation nDCG@10 (long-context + short adversarial) → epoch 6

Training curve

Framework versions

Python 3.10.14 · PyTorch 2.8.0+cu126 · Transformers 4.55.4 · Sentence Transformers 5.4.1 · Accelerate 1.10.0 · Tokenizers 0.21.0

Citation

@software{Mizan_Rerank_v3_2026,
  author    = {Ali Aljiachi},
  title     = {Mizan-Rerank-v3: Adversarially Trained Arabic Long-Context Reranker},
  year      = {2026},
  publisher = {Hugging Face},
  url       = {https://huggingface.co/ALJIACHI/Mizan-Rerank-v3}
}

License

Apache 2.0, the same license as the base model Alibaba-NLP/gte-multilingual-reranker-base.

Downloads last month
-
Safetensors
Model size
0.3B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for ALJIACHI/Mizan-Rerank-v3

Finetuned
(13)
this model

Dataset used to train ALJIACHI/Mizan-Rerank-v3

Space using ALJIACHI/Mizan-Rerank-v3 1

Collection including ALJIACHI/Mizan-Rerank-v3