quiz

A DistilBERT encoder fine-tuned for extractive question answering on SQuAD. Given a question and a passage, the model returns the span of text it identifies as the answer. It does not generate replies, and there is no conversation, retrieval or fallback logic anywhere in the repository. The weights are a single extractive span predictor.

The encoder is distilbert-base-uncased: six layers, 768 hidden dimensions, roughly 66M parameters, with a span-prediction head, stored in float32. Because the vocabulary is uncased WordPiece, input is lowercased and the model cannot distinguish Apple from apple. The 512-token position limit comes from DistilBERT and bounds the combined length of question and passage.

Usage

Transformers 5 removed the question-answering pipeline, so load the model directly:

import torch
from transformers import AutoModelForQuestionAnswering, AutoTokenizer

tok = AutoTokenizer.from_pretrained("harpertoken/quiz")
model = AutoModelForQuestionAnswering.from_pretrained("harpertoken/quiz")

question = "Who wrote Hamlet?"
context = "Hamlet is a tragedy written by William Shakespeare around 1600."
inputs = tok(question, context, return_tensors="pt", truncation=True, max_length=512)

with torch.inference_mode():
    out = model(**inputs)
start, end = int(out.start_logits.argmax()), int(out.end_logits.argmax())
print(tok.decode(inputs.input_ids[0][start : end + 1]))

On three check questions the model returns paris, william shakespeare and 1889. Answers come back lowercased, since the input was.

The repository now ships only model.safetensors; an earlier version also carried the same parameters as pytorch_model.bin, which doubled the download for no benefit.

Limitations

SQuAD is a reading-comprehension benchmark assembled from English Wikipedia paragraphs, and a model trained on it inherits that distribution. Performance on specialised domains, on questions requiring multi-hop reasoning, and on passages substantially longer than a Wikipedia paragraph will be lower than the in-domain numbers suggest, and I have not measured them here. Earlier versions of this card quoted exact-match and F1 scores and described TensorFlow support and multi-strategy response generation. The scores came from no evaluation I can point to, and no TensorFlow checkpoint was ever present, so both have been removed. If you need a number, evaluate on your own data with the SQuAD v1.1 dev set as a reference point.

Attribution

DistilBERT is described in Sanh et al., DistilBERT, a distilled version of BERT (2019). SQuAD is described in Rajpurkar et al., SQuAD: 100,000+ Questions for Machine Comprehension of Text (2016).

Downloads last month
252
Safetensors
Model size
66.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for harpertoken/quiz

Finetuned
(12618)
this model

Dataset used to train harpertoken/quiz