bert-base-squad-qa

Fine-tuning of bert-base-uncased for extractive Question Answering, delivered as part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science, Unit 2, Universidad Politécnica de Yucatán).

Model description

bert-base-uncased with a qa_outputs head that predicts, per token, the probability of being the start and the end of the answer span within the given context.

Training data

  • Dataset: SQuAD v1.1 (rajpurkar/squad), subsampled to 15,000 examples from the official train split (as required by the assignment), with a 90/10 split used for training/ validation. SQuAD's official validation set (10,570 questions, never seen during training) was reserved as the final test set.
  • Preprocessing uses a sliding window (stride) since contexts can exceed the maximum input length; each window is labeled with the start/end position of the answer inside that window, or (0, 0) if the answer does not fit.

Training procedure

Two adaptation methods were trained and compared; full fine-tuning is the delivered model, since it is necessary (not optional) for this task — partial fine-tuning underperforms by a wide margin.

Hyperparameter Value
Base model bert-base-uncased (110M params)
Method Full fine-tuning (BERT body + head, jointly)
Learning rate (head) 1e-3
Learning rate (BERT body) 2e-5
Epochs 3
Batch size 16 (train) / 32 (eval)
Seed 42
Trainable parameters 108,893,186
Training time 13.5 min (single T4 GPU)

Compared alternative (not delivered): partial fine-tuning, freezing the entire BERT body except its last 2 encoder layers (14,177,282 trainable params, 5.9 min).

Evaluation results

Metrics: official SQuAD Exact Match (EM) and word-level F1.

Method Val EM Val F1 Test EM Test F1
Partial fine-tuning (last 2 layers + head) 39.53% 55.05% 45.91% 59.45%
Full fine-tuning (delivered) 59.27% 73.79% 71.16% 81.06%

The 21.61-point F1 gap on test is by far the largest gap of the four tasks in this project (roughly 5x the gap seen in classification, NER, or POS), reflecting that extractive QA requires jointly reasoning over the question and locating a span in the context — something only 2 unfrozen layers cannot capture well.

Note on scale: this model was trained on 15k examples rather than the full 87.6k SQuAD train set, so its 81.06 test F1 is below the ~88 F1 reported in the literature for BERT-base trained on the full dataset — a result consistent with the reduced amount of training data.

Intended uses & limitations

  • Intended use: extractive question answering over short English passages, in the style of SQuAD v1.1, for coursework/research.
  • Limitations: trained on only 15k (of 87.6k) SQuAD examples, so accuracy is meaningfully below full-dataset BERT-base benchmarks. Only answers questions whose answer is a contiguous span present in the given context (no "no answer" case, as in SQuAD v1.1, and no closed-book/generative QA). Single seed, 3 epochs, no extensive hyperparameter search.

References

Downloads last month
25
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Dalila-Ku/bert-base-squad-qa

Finetuned
(7009)
this model

Dataset used to train Dalila-Ku/bert-base-squad-qa

Paper for Dalila-Ku/bert-base-squad-qa