SQuAD 1.1 extractive question answering

Selected method: full_finetuning, epoch 3. The choice was fixed on validation data before test evaluation. This is a measured academic adaptation of BERT, not a production-validated system.

Training data and evaluation

15,000 training / 2,000 internal validation / 2,000 held-out test questions. Reserve 15% of official training article titles for validation; test comes from official validation. No contexts overlap across selected subsets.

Normalized SQuAD 1.1 exact match and answer token F1 (0-100); validation F1 selects the checkpoint.

BERT predicts start/end positions. Max sequence 384, stride 128; preserve token_type_ids. Decode context-only top-20 start/end candidates, maximum answer length 30, best summed score across windows.

The table preserves native scales: AG News, NER and POS metrics are in [0, 1]; QA EM/F1 are in [0, 100]. Training seconds exclude validation, saving and final evaluation.

method best_epoch trainable_parameters train_seconds test_exact_match test_f1
partial_finetuning 3 14177282 305.736571 47.800000 61.749065
full_finetuning 3 108893186 845.982594 73.600000 83.277402

Optimization and reproducibility

Three epochs, seed 42. AdamW: head LR 1e-3, trainable encoder LR 2e-5, weight decay 0.01, 10% linear warmup, clipping 1.0. Effective batch size 16. A partial method trains only the last two encoder layers and head; lower layers are frozen in evaluation mode. Frozen weights were checked for invariance. Full fine-tuning updates all parameters.

Base model revision: 86b5e0934494bd15c9632b12f734a8a67f723594. Dataset revision: 7b6d24c440a36b6815f21b70d25016731768db1f. Detailed configuration is in training_config.json; the complete comparison is in evaluation.json. The included experiment source and requirements-lock.txt document the original environment. Use a new output directory when reproducing. QA source includes the documented UTF-8 JSON read fix; it did not change any training weights.

Only one seed was evaluated. Small gaps may reflect initialization, dropout and ordering variability; no statistical significance is claimed. Training hardware: RTX 4060 Ti, CUDA bf16. Training runtime is not inference latency.

Intended use and limitations

English answer extraction when an answer is present in a supplied context.

Does not generate new answers or evaluate abstention. Single seed and subset evaluation; SDPA backward can be nondeterministic. Not validated on Spanish or enterprise documents.

Loading the delivered model

from transformers import pipeline
qa = pipeline('question-answering', model='Terrificfantasm/bert-base-uncased-squad-qa')
print(qa(question='Where is the office?', context='The office is in London.',
         max_seq_len=384, doc_stride=128, max_answer_len=30))

This is an illustrative loading example. Exact benchmark reproduction uses the included experiment's cross-window decoder, not the pipeline's potentially different postprocessing.

Source terms and references

The upstream BERT checkpoints identify Apache-2.0 licensing. The SQuAD distribution is identified as CC BY-SA 4.0. Dataset terms are separate from the base checkpoint license. This repository does not redistribute the training corpus. It preserves upstream attribution without asserting a new blanket license over all data sources.

Downloads last month
18
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Terrificfantasm/bert-base-uncased-squad-qa

Finetuned
(6996)
this model

Dataset used to train Terrificfantasm/bert-base-uncased-squad-qa

Paper for Terrificfantasm/bert-base-uncased-squad-qa