BERT (base-cased) adapted for four NLP tasks

This repo contains four independently fine-tuned bert-base-cased models, one per task, each in its own subfolder. All were trained for a course assignment comparing feature-based, partial, and full fine-tuning adaptation methods.

Subfolder Task Dataset Adaptation method
topic_classification/ Topic classification (4-way) AG News Full fine-tuning
ner/ Named Entity Recognition CoNLL-2003 Full fine-tuning
pos/ POS Tagging (17-tag UPOS) UD English EWT Full fine-tuning
qa/ Extractive QA SQuAD v1.1 Full fine-tuning

How to load a model

from transformers import AutoModelForSequenceClassification, AutoTokenizer

model = AutoModelForSequenceClassification.from_pretrained(
    "Dodi488/BERT_for_NLP", subfolder="topic_classification"
)
tokenizer = AutoTokenizer.from_pretrained(
    "Dodi488/BERT_for_NLP", subfolder="topic_classification"
)

Swap AutoModelForSequenceClassification for AutoModelForTokenClassification (ner, pos) or AutoModelForQuestionAnswering (qa), and the subfolder name accordingly.


Topic Classification (topic_classification/)

4-class AG News topic classification (World, Sports, Business, Sci/Tech). Trained on a 2,000-example subset (500 eval), full fine-tuning with split learning rate (head 1e-3, encoder 2e-5), compared against a frozen-feature Logistic Regression baseline.

Method Accuracy Macro F1
Feature-based (LogReg) 0.8080 0.8021
Full fine-tuning (delivered) 0.8740 0.8687

Named Entity Recognition (ner/)

CoNLL-2003 NER (PER/ORG/LOC/MISC, BIO-tagged). Trained on a 2,000-example subset (500 eval), full fine-tuning delivered over partial fine-tuning (top 2 encoder layers + head).

Method Eval Loss Accuracy F1
Partial fine-tuning 0.2113 0.9493 0.7849
Full fine-tuning (delivered) 0.0876 0.9779 0.8873

Full fine-tuning shows a ~10.2-point F1 gain over partial, a larger gap than initially hypothesized for a token-level task -- NER benefits from adapting more of the encoder than just the top layers.

POS Tagging (pos/)

UD English EWT, 17-tag Universal POS. Trained on a 2,000-example subset (500 eval), full fine-tuning delivered over partial fine-tuning.

Method Eval Loss Accuracy F1
Partial fine-tuning 0.3043 0.9181 0.9038
Full fine-tuning (delivered) 0.1497 0.9610 0.9541

Partial fine-tuning recovers most of full fine-tuning's quality (~5-point F1 gap) at 1/8th the trainable parameters -- a reasonable cheaper choice if compute is constrained.

Extractive QA (qa/)

SQuAD v1.1 span prediction. Trained on a 15,000-example subset (1,000 eval), full fine-tuning delivered over partial fine-tuning.

Method Eval Loss Exact Match F1
Partial fine-tuning 1.8480 0.3770 0.5031
Full fine-tuning (delivered) 1.3528 0.5770 0.7066

Largest method gap of any task in this project (20 points EM/F1), confirming QA needs more adaptation capacity than tagging or classification. Note: EM/F1 are computed over predicted vs. gold token-index spans rather than decoded answer text -- a lighter-weight proxy for the official evaluate.load("squad") script, expected to track it closely but not match it exactly. Full-fine-tuning F1 (0.7066) is below published BERT-base SQuAD v1.1 benchmarks (88 F1), likely reflecting the 2-epoch training budget alongside this scoring convention.


General limitations (all four models)

  • All classification/tagging tasks trained on 2,000-example subsets, not full dataset splits; numbers may shift with more data.
  • One run per configuration, no seed-variance estimate (course guidance: ±1-3 points expected noise from seed/init/dropout alone).
  • Not validated outside the training domain/language/label set for each task.

References

  • Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. https://arxiv.org/abs/1810.04805
  • Rajpurkar, P. et al. (2016). SQuAD: 100,000+ Questions for Machine Comprehension of Text. https://arxiv.org/abs/1606.05250
  • Datasets: fancyzhx/ag_news, lhoestq/conll2003, universal_dependencies (en_ewt), rajpurkar/squad -- all via Hugging Face datasets.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Dodi488/BERT_for_NLP

Finetuned
(2944)
this model

Datasets used to train Dodi488/BERT_for_NLP

Papers for Dodi488/BERT_for_NLP