Instructions to use toyin88/u2t01-bert-qa with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use toyin88/u2t01-bert-qa with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("question-answering", model="toyin88/u2t01-bert-qa")# Load model directly from transformers import AutoTokenizer, AutoModelForQuestionAnswering tokenizer = AutoTokenizer.from_pretrained("toyin88/u2t01-bert-qa") model = AutoModelForQuestionAnswering.from_pretrained("toyin88/u2t01-bert-qa", device_map="auto") - Notebooks
- Google Colab
- Kaggle
SQuAD v1.1 extractive question answering — full adaptation
Part of U2T01: Adapting BERT for NLP tasks. One of four models, each one a
different head on the same bert-base-uncased body, each delivered with
the adaptation method that its own measurements justified.
What this model does
Predicts start and end token positions of the answer span.
How it was adapted
Full fine-tuning. Every encoder layer and the head were trained. The embedding matrix was kept frozen (23M of BERT's 110M parameters), which changes nothing measurable at these dataset sizes and saves optimiser memory.
| Base model | bert-base-uncased |
| Adaptation | full |
| Trainable parameters | 85,056,002 of 108,893,186 (78.11%) |
| Encoder layers trained | 12 of 12 |
| Head learning rate | 0.001 |
| Encoder learning rate | 3e-05 |
| Epochs | 1 |
| Batch size | 16 |
| Max sequence length | 384 |
| Seed | 42 |
| Training time | 1.0 min on Tesla T4 |
A freshly initialised head and pretrained encoder layers were given separate learning rates through two optimiser parameter groups; a single learning rate for both silently cripples this kind of run.
Training data
SQuAD v1.1 (rajpurkar/squad).
Training examples used: 2000.
Sub-sampled to ~15k training examples as the assignment specifies.
Evaluation
| Metric | Test | Validation |
|---|---|---|
| exact_match | 37.30 | 29.50 |
| f1 | 47.74 | 43.68 |
Metric choice is argued in the project report. In short: entity-level F1 for
NER because ~83% of CoNLL tokens are O and token accuracy would flatter a
model that predicts nothing; token accuracy for POS because there is no
dominant class; EM and F1 for SQuAD because each covers the other's blind
spot.
Why this method, and what lost
| Run | Method | Trainable params | f1 (test) |
|---|---|---|---|
qa_full_mini |
full | 85,056,002 | 47.74 |
qa_partial_mini |
partial | 28,353,026 | 30.78 |
Decision rule: best primary metric on test, and where two methods fall inside a ±1 point band — smaller than the ±1–3 points a different random seed moves results by — the cheaper one wins.
Intended use and limitations
Intended for research and coursework: benchmarking adaptation strategies and as a starting point for related English tasks.
Not intended for production decisions about people. Concretely:
- English only, and the training text is narrow. AG News is 2004 newswire; CoNLL-2003 is Reuters newswire, so it under-performs on social media, names outside Western European conventions, and any entity that became prominent after 2003. UD EWT is web English (blogs, reviews, emails). SQuAD contexts are Wikipedia paragraphs, and the model always returns a span — SQuAD v1.1 contains no unanswerable questions, so it cannot say "I don't know".
- Single run, single seed. Re-running with a different seed moves results by roughly ±1–3 points. Differences smaller than that are not meaningful.
- Inherited bias.
bert-base-uncasedcarries the associations of its pretraining corpus, and nothing here mitigates them. - Sub-sampled training data, so these numbers sit below published full-data results. Every adaptation method on this task saw exactly the same data, so the comparison that chose this method is unaffected.
Reproduce it
git clone <this project>
pip install -r requirements.txt
python -m u2t01.cli run configs_mini/qa_full.yaml
Fixed seed (42), pinned dependencies, and one YAML file per experiment: config + seed + requirements.txt fully determine a run.
References
- Devlin et al., BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding, 2018 — https://arxiv.org/abs/1810.04805 (§5.3 is the feature-based vs fine-tuning comparison this project mirrors)
- Tunstall, von Werra & Wolf, Natural Language Processing with Transformers, O'Reilly, chapters 1–3
- Hugging Face, Fine-tune a pretrained model — https://huggingface.co/docs/transformers/training
- Dataset:
rajpurkar/squad
Card generated 2026-09-27 from results/qa_full_mini.json.
- Downloads last month
- 15
Model tree for toyin88/u2t01-bert-qa
Base model
google-bert/bert-base-uncased