BERT (base-cased) adapted for four NLP tasks
This repo contains four independently fine-tuned bert-base-cased models,
one per task, each in its own subfolder. All were trained for a course
assignment comparing feature-based, partial, and full fine-tuning
adaptation methods.
| Subfolder | Task | Dataset | Adaptation method |
|---|---|---|---|
topic_classification/ |
Topic classification (4-way) | AG News | Full fine-tuning |
ner/ |
Named Entity Recognition | CoNLL-2003 | Full fine-tuning |
pos/ |
POS Tagging (17-tag UPOS) | UD English EWT | Full fine-tuning |
qa/ |
Extractive QA | SQuAD v1.1 | Full fine-tuning |
How to load a model
from transformers import AutoModelForSequenceClassification, AutoTokenizer
model = AutoModelForSequenceClassification.from_pretrained(
"Dodi488/BERT_for_NLP", subfolder="topic_classification"
)
tokenizer = AutoTokenizer.from_pretrained(
"Dodi488/BERT_for_NLP", subfolder="topic_classification"
)
Swap AutoModelForSequenceClassification for AutoModelForTokenClassification
(ner, pos) or AutoModelForQuestionAnswering (qa), and the subfolder name
accordingly.
Topic Classification (topic_classification/)
4-class AG News topic classification (World, Sports, Business, Sci/Tech). Trained on a 2,000-example subset (500 eval), full fine-tuning with split learning rate (head 1e-3, encoder 2e-5), compared against a frozen-feature Logistic Regression baseline.
| Method | Accuracy | Macro F1 |
|---|---|---|
| Feature-based (LogReg) | 0.8080 | 0.8021 |
| Full fine-tuning (delivered) | 0.8740 | 0.8687 |
Named Entity Recognition (ner/)
CoNLL-2003 NER (PER/ORG/LOC/MISC, BIO-tagged). Trained on a 2,000-example subset (500 eval), full fine-tuning delivered over partial fine-tuning (top 2 encoder layers + head).
| Method | Eval Loss | Accuracy | F1 |
|---|---|---|---|
| Partial fine-tuning | 0.2113 | 0.9493 | 0.7849 |
| Full fine-tuning (delivered) | 0.0876 | 0.9779 | 0.8873 |
Full fine-tuning shows a ~10.2-point F1 gain over partial, a larger gap than initially hypothesized for a token-level task -- NER benefits from adapting more of the encoder than just the top layers.
POS Tagging (pos/)
UD English EWT, 17-tag Universal POS. Trained on a 2,000-example subset (500 eval), full fine-tuning delivered over partial fine-tuning.
| Method | Eval Loss | Accuracy | F1 |
|---|---|---|---|
| Partial fine-tuning | 0.3043 | 0.9181 | 0.9038 |
| Full fine-tuning (delivered) | 0.1497 | 0.9610 | 0.9541 |
Partial fine-tuning recovers most of full fine-tuning's quality (~5-point F1 gap) at 1/8th the trainable parameters -- a reasonable cheaper choice if compute is constrained.
Extractive QA (qa/)
SQuAD v1.1 span prediction. Trained on a 15,000-example subset (1,000 eval), full fine-tuning delivered over partial fine-tuning.
| Method | Eval Loss | Exact Match | F1 |
|---|---|---|---|
| Partial fine-tuning | 1.8480 | 0.3770 | 0.5031 |
| Full fine-tuning (delivered) | 1.3528 | 0.5770 | 0.7066 |
Largest method gap of any task in this project (20 points EM/F1),
confirming QA needs more adaptation capacity than tagging or classification.
Note: EM/F1 are computed over predicted vs. gold token-index spans
rather than decoded answer text -- a lighter-weight proxy for the official
88 F1), likely reflecting the 2-epoch
training budget alongside this scoring convention.evaluate.load("squad") script, expected to track it closely but not
match it exactly. Full-fine-tuning F1 (0.7066) is below published
BERT-base SQuAD v1.1 benchmarks (
General limitations (all four models)
- All classification/tagging tasks trained on 2,000-example subsets, not full dataset splits; numbers may shift with more data.
- One run per configuration, no seed-variance estimate (course guidance: ±1-3 points expected noise from seed/init/dropout alone).
- Not validated outside the training domain/language/label set for each task.
References
- Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. https://arxiv.org/abs/1810.04805
- Rajpurkar, P. et al. (2016). SQuAD: 100,000+ Questions for Machine Comprehension of Text. https://arxiv.org/abs/1606.05250
- Datasets: fancyzhx/ag_news, lhoestq/conll2003, universal_dependencies
(en_ewt), rajpurkar/squad -- all via Hugging Face
datasets.
Model tree for Dodi488/BERT_for_NLP
Base model
google-bert/bert-base-cased