Question Answering
Transformers
Safetensors
English
bert
qa
full
u2t01

BERT adapted to extractive question answering (SQuAD v1.1) — full fine-tuning

abelcetina/u2t01-bert-squad — bert-base-uncased adapted to extractive question answering (SQuAD v1.1) with full fine-tuning.

Produced for the university assignment U2T01: Adapting BERT for NLP tasks. The companion repository contains the training code, the raw results file behind every number below, and the comparison against the alternative adaptation methods.

Model description

  • Base encoder: bert-base-uncased (108,893,186 parameters in total).
  • Task: extractive question answering (SQuAD v1.1).
  • Adaptation method: full fine-tuning. Every parameter of the encoder and of the task head received gradients, with a small learning rate on the body and a larger one on the randomly initialised head.
  • Head: two linear outputs predicting the start and the end token of the answer span.
  • Label set: start and end token positions of the answer span inside the context
  • Trained by: U2T01 team · generated 2026-09-26.

Intended use

  • Research and teaching: reproducing the feature-based versus fine-tuning comparison of BERT section 5.3 across four task types.
  • Inference on English text that resembles the training domain (paragraphs from English Wikipedia with crowd-written questions).
  • Usage:
from transformers import pipeline
qa = pipeline("question-answering", model="abelcetina/u2t01-bert-squad")
print(qa(question="Where was the treaty signed?",
         context="The treaty was signed in Lisbon in December 2007."))

Out-of-scope use

  • Any decision affecting a person (hiring, moderation with consequences, credit, legal or medical decisions). This is a coursework model validated on one academic benchmark with a single seed.
  • Languages other than English, and domains far from paragraphs from English Wikipedia with crowd-written questions — performance is not characterised there and is expected to drop.
  • Extracting facts about the world: for the QA model, the answer is a span copied from the context passage it is given, never knowledge the model holds.
  • Commercial redistribution without checking the training-data licence below.

Training data

Dataset SQuAD v1.1
Identifier used in code rajpurkar/squad
Licence CC BY-SA 4.0
Splits train 87,599 · validation 10,570 questions (a subsample was used for training)
Training examples actually used 15,000
Evaluation split reported below validation

Around 15,000 training examples were subsampled to fit the Colab T4 budget, so absolute EM/F1 are below a full-data run.

Adaptation method and what was frozen

  • Method flag: --method full
  • Trainable components: the whole BERT encoder (embeddings + 12 layers) and the task head
  • Frozen components: nothing
Parameters Value
Total 108,893,186
Trainable 108,893,186
Trainable share 100 %

Hyper-parameters

Hyper-parameter Value
Epochs 2
Batch size 16
Head learning rate 0.001
Body learning rate 0.00003
Weight decay 0.01
Max sequence length 384
Top layers unfrozen all 12 encoder layers (full fine-tuning)
Seed 42

Evaluation results

Metrics computed by src/metrics.py; the raw values live in qa__full__seed42.json.

Evaluation split (validation)

Metric Value
exact_match 72.5639
f1 81.7763

Test split

This run was not evaluated on a held-out test split.

Headline number: f1 = 81.7763 (exact_match = 72.5639).

The comparison against the other adaptation methods of this task, with the gap in percentage points and the noise flag, is in report/results_tables.md of the companion code repository (Table 1, Table 2.qa and Table 3). A gap below one point there is reported as noise, not as an improvement.

Limitations and bias

  • Dataset bias. SQuAD v1.1 questions are crowd-written while looking at the paragraph, so they share vocabulary with the answer far more than real questions do, and every question is answerable — the model has no way to abstain, and will always return a span. The model inherits both that domain skew and the social biases of the bert-base-uncased pre-training corpus (BookCorpus and English Wikipedia), which is known to encode gender, ethnic and occupational stereotypes. Nothing here mitigates or measures them.
  • Single seed. Every number on this card comes from one run with seed 42. No variance estimate exists, so a difference of roughly one point against another configuration must be read as noise rather than as an improvement.
  • Subsampling. Training used 15,000 examples (hyperparams.train_size). Around 15,000 training examples were subsampled to fit the Colab T4 budget, so absolute EM/F1 are below a full-data run.
  • Sequence truncation. Inputs were truncated to 384 tokens; longer inputs lose their tail at inference too.
  • Frozen-encoder caveat (feature-based models only). Not applicable: this model was fine-tuned, so the encoder weights moved.
  • Only a validation (and where available test) split was measured. No robustness, adversarial, out-of-domain or fairness evaluation was performed.

Environmental and compute footprint

Hardware 1 × Tesla T4 (Google Colab), CUDA 12.8
Training wall time 638.1 s (~10.6 min) over 1,894 steps
Seconds per step 0.337
Peak GPU memory 3905 MB
Software transformers 5.16.1 · torch 2.11.0+cu128 · datasets 4.8.5
Run timestamp 2026-09-25T23:56:04

A single short fine-tuning run on one T4 is a small footprint in absolute terms, and that is the point of this assignment: the cost-versus-quality table in the companion report shows how much of the quality survives when only a fraction of the parameters is trained, which is the part that scales. No CO₂e figure is claimed here, because the energy mix of the Colab region is unknown.

References and citation

  • Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL-HLT. arXiv:1810.04805. Section 5.3 is the feature-based versus fine-tuning comparison this model is part of.
  • Rajpurkar, P., Zhang, J., Lopyrev, K., & Liang, P. (2016). SQuAD: 100,000+ Questions for Machine Comprehension of Text. EMNLP. arXiv:1606.05250.
  • Tunstall, L., von Werra, L., & Wolf, T. (2022). Natural Language Processing with Transformers, chapters 1–3. O'Reilly.
  • Hugging Face, Fine-tuning a pretrained model: https://huggingface.co/docs/transformers/training
@inproceedings{devlin2019bert,
  title     = {BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding},
  author    = {Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina},
  booktitle = {Proceedings of NAACL-HLT},
  pages     = {4171--4186},
  year      = {2019},
  url       = {https://arxiv.org/abs/1810.04805}
}

Cite this adapted model as:

@misc{u2t01_qa_full,
  title  = {BERT adapted to extractive question answering (SQuAD v1.1) — full fine-tuning},
  author = {U2T01 team},
  year   = {2026},
  note   = {U2T01: Adapting BERT for NLP tasks. Adaptation method: full. Seed 42.},
  url    = {https://huggingface.co/abelcetina/u2t01-bert-squad}
}
Downloads last month
16
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for abelcetina/u2t01-bert-squad

Finetuned
(6996)
this model

Dataset used to train abelcetina/u2t01-bert-squad

Papers for abelcetina/u2t01-bert-squad