clue

A short continued-fine-tuning run of quiz, itself a DistilBERT encoder adapted for extractive question answering on SQuAD. The architecture and tokenizer are identical; only the weights differ. Where quiz reflects a full training pass, this checkpoint reflects roughly a thousand SQuAD examples seen once, which makes it a useful small-scale reference point and a poor substitute for a properly trained model.

Training used a learning rate of 2e-5 at batch size one for a single epoch, in float32. The published weights are model.safetensors. The config.json previously carried a key tie_weights_, which no version of Transformers reads; it has been removed, and nothing else in the config was altered.

Usage

Transformers 5 removed the question-answering pipeline, so load the model directly:

import torch
from transformers import AutoModelForQuestionAnswering, AutoTokenizer

tok = AutoTokenizer.from_pretrained("harpertoken/clue")
model = AutoModelForQuestionAnswering.from_pretrained("harpertoken/clue")

question = "Who wrote Hamlet?"
context = "Hamlet is a tragedy written by William Shakespeare around 1600."
inputs = tok(question, context, return_tensors="pt", truncation=True, max_length=512)

with torch.inference_mode():
    out = model(**inputs)
start, end = int(out.start_logits.argmax()), int(out.end_logits.argmax())
print(tok.decode(inputs.input_ids[0][start : end + 1]))

On the three questions used to check quiz (the capital of France, the author of Hamlet, and the year the Eiffel Tower was completed), this checkpoint returns paris, william shakespeare and 1889, the same answers. A thousand examples has not visibly degraded it, which is itself a reason to doubt that the fine-tuning taught much.

Limitations

A thousand examples is a demonstration of the fine-tuning mechanics rather than a trained model, and the documentation this replaced claimed SQuAD exact-match and F1 figures that were never produced by an evaluation. Treat this as a low-fidelity copy of quiz. It is English-only, inherits the same uncased tokenisation and SQuAD domain bias described in the quiz card, and shares its 512-token limit. Compare the two directly before assuming the fine-tuning helped.

Attribution

DistilBERT follows Sanh et al. (2019); SQuAD follows Rajpurkar et al. (2016).

Downloads last month
29
Safetensors
Model size
66.4M params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for harpertoken/clue

Finetuned
(12581)
this model

Dataset used to train harpertoken/clue