bert-base-squad-qa
Fine-tuning of bert-base-uncased for extractive Question Answering, delivered as
part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science,
Unit 2, Universidad Politécnica de Yucatán).
Model description
bert-base-uncased with a qa_outputs head that predicts, per token, the
probability of being the start and the end of the answer span within the given
context.
Training data
- Dataset: SQuAD v1.1
(
rajpurkar/squad), subsampled to 15,000 examples from the official train split (as required by the assignment), with a 90/10 split used for training/ validation. SQuAD's official validation set (10,570 questions, never seen during training) was reserved as the final test set. - Preprocessing uses a sliding window (stride) since contexts can exceed the maximum
input length; each window is labeled with the start/end position of the answer
inside that window, or
(0, 0)if the answer does not fit.
Training procedure
Two adaptation methods were trained and compared; full fine-tuning is the delivered model, since it is necessary (not optional) for this task — partial fine-tuning underperforms by a wide margin.
| Hyperparameter | Value |
|---|---|
| Base model | bert-base-uncased (110M params) |
| Method | Full fine-tuning (BERT body + head, jointly) |
| Learning rate (head) | 1e-3 |
| Learning rate (BERT body) | 2e-5 |
| Epochs | 3 |
| Batch size | 16 (train) / 32 (eval) |
| Seed | 42 |
| Trainable parameters | 108,893,186 |
| Training time | 13.5 min (single T4 GPU) |
Compared alternative (not delivered): partial fine-tuning, freezing the entire BERT body except its last 2 encoder layers (14,177,282 trainable params, 5.9 min).
Evaluation results
Metrics: official SQuAD Exact Match (EM) and word-level F1.
| Method | Val EM | Val F1 | Test EM | Test F1 |
|---|---|---|---|---|
| Partial fine-tuning (last 2 layers + head) | 39.53% | 55.05% | 45.91% | 59.45% |
| Full fine-tuning (delivered) | 59.27% | 73.79% | 71.16% | 81.06% |
The 21.61-point F1 gap on test is by far the largest gap of the four tasks in this project (roughly 5x the gap seen in classification, NER, or POS), reflecting that extractive QA requires jointly reasoning over the question and locating a span in the context — something only 2 unfrozen layers cannot capture well.
Note on scale: this model was trained on 15k examples rather than the full 87.6k SQuAD train set, so its 81.06 test F1 is below the ~88 F1 reported in the literature for BERT-base trained on the full dataset — a result consistent with the reduced amount of training data.
Intended uses & limitations
- Intended use: extractive question answering over short English passages, in the style of SQuAD v1.1, for coursework/research.
- Limitations: trained on only 15k (of 87.6k) SQuAD examples, so accuracy is meaningfully below full-dataset BERT-base benchmarks. Only answers questions whose answer is a contiguous span present in the given context (no "no answer" case, as in SQuAD v1.1, and no closed-book/generative QA). Single seed, 3 epochs, no extensive hyperparameter search.
References
- Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805. https://arxiv.org/abs/1810.04805
- Rajpurkar, P., Zhang, J., Lopyrev, K., & Liang, P. (2016). SQuAD: 100,000+ Questions for Machine Comprehension of Text.
- Hugging Face. Fine-tune a pretrained model. https://huggingface.co/docs/transformers/training
- Dataset: rajpurkar/squad
- Downloads last month
- 25
Model tree for Dalila-Ku/bert-base-squad-qa
Base model
google-bert/bert-base-uncased