DISTO

DISTO is a learned metric for the quality of distractors in multiple-choice reading comprehension questions. Given an article, a question, the correct answer and up to three distractors, it returns a score in [0, 1]: higher means the distractors are more plausible in context. It needs no reference distractors.

How to run

Option 1: the disto package (recommended)

pip install git+https://github.com/bilalghanem/DISTO.git
from disto import DistoScorer

scorer = DistoScorer("bilalghanem/DISTO")  # downloads the model from the Hub

article = "The Nile is the longest river in Africa. It flows north through Egypt into the Mediterranean Sea."
question = "Where does the Nile flow into?"
answer = "The Mediterranean Sea"

# Score the whole set of up to three distractors in one pass
print(scorer.score(article, question, answer, ["The Red Sea", "The Atlantic Ocean", "Lake Victoria"]))  # ~0.99

# A distractor that copies the answer or is unrelated scores low
print(scorer.score(article, question, answer, ["The Mediterranean Sea", "A bowl of soup", "Tuesday"]))  # ~0.01

# Score a single distractor
print(scorer.score(article, question, answer, ["The Red Sea"]))  # ~0.997

# Score each distractor on its own and average, as in the paper
scores = scorer.score_each(article, question, answer, ["The Red Sea", "The Atlantic Ocean", "Lake Victoria"])
print(sum(scores) / len(scores))

Option 2: plain transformers

import torch
from transformers import AutoModelForSequenceClassification, AutoTokenizer

tokenizer = AutoTokenizer.from_pretrained("bilalghanem/DISTO")
model = AutoModelForSequenceClassification.from_pretrained("bilalghanem/DISTO").eval()

text = (
    "[QUES] Where does the Nile flow into? "
    "[ANS] The Mediterranean Sea "
    "[DIS1] The Red Sea [DIS2] The Atlantic Ocean [DIS3] Lake Victoria "
    "[ART] The Nile is the longest river in Africa. It flows north through Egypt into the Mediterranean Sea."
)
inputs = tokenizer(text, truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
    score = torch.sigmoid(model(**inputs).logits).item()  # the sigmoid is not part of the model
print(score)

To score a single distractor with plain transformers, fill the other two slots with [EMPT]:

text = (
    "[QUES] Where does the Nile flow into? "
    "[ANS] The Mediterranean Sea "
    "[DIS1] The Red Sea [DIS2] [EMPT] [DIS3] [EMPT] "
    "[ART] The Nile is the longest river in Africa. It flows north through Egypt into the Mediterranean Sea."
)
inputs = tokenizer(text, truncation=True, max_length=512, return_tensors="pt")
with torch.no_grad():
    print(torch.sigmoid(model(**inputs).logits).item())  # ~0.997

Input format

The input is [QUES] question [ANS] answer [DIS1] d1 [DIS2] d2 [DIS3] d3 [ART] article. Fill missing distractors with [EMPT] (the paper's single-distractor setting is one distractor in [DIS1] and [EMPT] in the other two). Put the article last, since only the article is cut at 512 tokens. The Hub inference widget does not build this format or apply the sigmoid, so use one of the options above.

Model details

  • Architecture: distilroberta-base with a single-output regression head (RobertaForSequenceClassification, num_labels=1) and a sigmoid applied to the logit. Seven special tokens are added to the tokenizer.
  • Max input length: 512 tokens (the article comes last and is the only part truncated).
  • Training: MSE loss, AdamW, learning rate 3e-5, batch size 20, early stopping on validation loss.
  • Training data: CosmosQA, DREAM, MCScript, MCTest, QuAIL, RACE and SciQ. Good distractors have a target of 1; negative samples are built by copying the answer, drawing a random distractor, taking the farthest point of the distractor's k-means cluster, or rewriting a distractor with BERT [MASK] filling.

Evaluation

Held-out test split of the release (66,322 instances): MAE 0.0286, Pearson correlation 0.966.

In the paper, DISTO correlates with Amazon Mechanical Turk ratings of distractor quality at Pearson 0.81 on gold and negatively sampled distractors, and between 0.28 and 0.75 on distractors produced by five distractor generation models.

Limitations

  • The score reflects consistency with the context, learned from English reading-comprehension datasets. It is not a replacement for review by an educator.
  • Human agreement on this task is only moderate (Fleiss' kappa 0.45 in the paper's study).
  • Only English has been tested, and only up to three distractors per question.

Citation

@inproceedings{ghanem2024disto,
  title     = {{DISTO}: Textual Distractors for Multiple Choice Reading Comprehension Questions Using Negative Sampling},
  author    = {Ghanem, Bilal and Fyshe, Alona},
  booktitle = {Proceedings of the 17th International Conference on Educational Data Mining},
  pages     = {6--17},
  year      = {2024},
  doi       = {10.5281/zenodo.12729766}
}
Downloads last month
27
Safetensors
Model size
82.1M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for bilalghanem/DISTO

Finetuned
(786)
this model

Datasets used to train bilalghanem/DISTO