halluscoring-camelbert-nli

CAMeLBERT (CAMeL-Lab/bert-base-arabic-camelbert-mix) fine-tuned on HalluScoring 2026 Task 1.1 using an NLI framing: [CLS] gold_answer [SEP] model_answer [SEP], treating hallucination detection as "does the model's answer entail/contradict the gold answer." Internally this is run S02 — the experiment that established NLI framing as the single biggest lever in this task (+5.6pp clean-dev AUC-ROC over the QA-framed halluscoring-camelbert-qa). Every model after this one uses the same NLI framing.

Not submitted to the competition — kept as an internal experiment and as a component of the S23/S24v/S25v ensembles. See SYSTEM_WRITEUP.md for the officially-submitted models.

How to Use

import torch
from transformers import AutoTokenizer, AutoModelForSequenceClassification

model_id = "HassanB4/halluscoring-camelbert-nli"
tokenizer = AutoTokenizer.from_pretrained(model_id)
model = AutoModelForSequenceClassification.from_pretrained(model_id)
model.eval()

gold_answer = "..."
model_answer = "..."

inputs = tokenizer(gold_answer, model_answer, truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
    logits = model(**inputs).logits
    prob_hallucinated = torch.softmax(logits, dim=-1)[0, 1].item()

print(f"hallucinated={int(prob_hallucinated > 0.5)}, score={prob_hallucinated:.4f}")

Training

Parameter Value
Base model CAMeL-Lab/bert-base-arabic-camelbert-mix
Input format nli (gold_answer + model_answer)
Max sequence length 512
Batch size 16
Epochs 5
Learning rate 2e-5
Warmup ratio 0.1
Weight decay 0.01
Loss cross-entropy
Seed 42

Evaluation

Split AUC-ROC F1-Macro
Dev (official, n=1300) 0.9574 0.9081
Dev (clean, unseen-question subset, n=800) 0.9272

Clean-dev AUC-ROC is our internal generalization estimate, restricted to the ~800 dev rows whose questions never appear in training (see SYSTEM_WRITEUP.md §"Key finding" for why).

Limitations

Not evaluated on the hidden test set — internal dev-only experiment. Used as a component of the S23, S24v, and S25v soft-vote ensembles.

Citation

@inproceedings{namaa2026halluscoring,
    title={{NAMAA at HalluScoring 2026: NLI-Framed BERT Classifiers and Ensembling for Model-Agnostic Arabic Hallucination Detection}},
    author={[AUTHOR NAMES TBD]},
    year={2026},
    booktitle={Proceedings of ArabicNLP 2026},
    note={HalluScoring 2026 Shared Task, Track 1}
}
Downloads last month
2
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for HassanB4/halluscoring-camelbert-nli

Finetuned
(11)
this model

Collection including HassanB4/halluscoring-camelbert-nli

Evaluation results

  • Clean Dev AUC-ROC (unseen questions) on HalluScoring 2026 Track 1, Task 1.1
    self-reported
    0.927
  • Official Dev AUC-ROC on HalluScoring 2026 Track 1, Task 1.1
    self-reported
    0.957
  • Official Dev F1-Macro on HalluScoring 2026 Track 1, Task 1.1
    self-reported
    0.908