BERT adapted to named entity recognition (CoNLL-2003, IOB2) — full fine-tuning

abelcetina/u2t01-bert-conll2003-ner — bert-base-uncased adapted to named entity recognition (CoNLL-2003, IOB2) with full fine-tuning.

Produced for the university assignment U2T01: Adapting BERT for NLP tasks. The companion repository contains the training code, the raw results file behind every number below, and the comparison against the alternative adaptation methods.

Model description

  • Base encoder: bert-base-uncased (108,898,569 parameters in total).
  • Task: named entity recognition (CoNLL-2003, IOB2).
  • Adaptation method: full fine-tuning. Every parameter of the encoder and of the task head received gradients, with a small learning rate on the body and a larger one on the randomly initialised head.
  • Head: a linear token-classification head over every token representation.
  • Label set: 9 IOB2 tags: O, B-PER, I-PER, B-ORG, I-ORG, B-LOC, I-LOC, B-MISC, I-MISC
  • Trained by: U2T01 team · generated 2026-09-25.

Intended use

  • Research and teaching: reproducing the feature-based versus fine-tuning comparison of BERT section 5.3 across four task types.
  • Inference on English text that resembles the training domain (Reuters newswire text from 1996–1997).
  • Usage:
from transformers import pipeline
tagger = pipeline("token-classification", model="abelcetina/u2t01-bert-conll2003-ner", aggregation_strategy="simple")
print(tagger("Angela Merkel visited the European Central Bank in Frankfurt."))

Out-of-scope use

  • Any decision affecting a person (hiring, moderation with consequences, credit, legal or medical decisions). This is a coursework model validated on one academic benchmark with a single seed.
  • Languages other than English, and domains far from Reuters newswire text from 1996–1997 — performance is not characterised there and is expected to drop.
  • Extracting facts about the world: for the QA model, the answer is a span copied from the context passage it is given, never knowledge the model holds.
  • Commercial redistribution without checking the training-data licence below.

Training data

Dataset CoNLL-2003 (parquet mirror)
Identifier used in code lhoestq/conll2003
Licence other — derived from the Reuters-21578 news corpus; research use under the Reuters terms referenced on the dataset card
Splits train 14,041 · validation 3,250 · test 3,453 sentences
Training examples actually used 14,041
Evaluation split reported below validation

Read from a parquet mirror rather than from the legacy script-based conll2003, because datasets >= 4.0 removed dataset loading scripts.

Adaptation method and what was frozen

  • Method flag: --method full
  • Trainable components: the whole BERT encoder (embeddings + 12 layers) and the task head
  • Frozen components: nothing
Parameters Value
Total 108,898,569
Trainable 108,898,569
Trainable share 100 %

Hyper-parameters

Hyper-parameter Value
Epochs 3
Batch size 32
Head learning rate 0.001
Body learning rate 0.00002
Weight decay 0.01
Max sequence length 128
Top layers unfrozen all 12 encoder layers (full fine-tuning)
Seed 42

Evaluation results

Metrics computed by src/metrics.py; the raw values live in ner__full__seed42.json.

Evaluation split (validation)

Metric Value
loss 0.0462
precision 0.9338
recall 0.9453
f1 0.9395
accuracy 0.9882

Test split

Metric Value
loss 0.1009
precision 0.8937
recall 0.9102
f1 0.9019
accuracy 0.9805

Headline number: f1 = 0.9395 (accuracy = 0.9882).

The comparison against the other adaptation methods of this task, with the gap in percentage points and the noise flag, is in report/results_tables.md of the companion code repository (Table 1, Table 2.ner and Table 3). A gap below one point there is reported as noise, not as an improvement.

Limitations and bias

  • Dataset bias. Reuters newswire from the late 1990s: person, organisation and location names are heavily skewed towards the actors and the spelling conventions of that period and towards sports reporting, which dominates the corpus. The model inherits both that domain skew and the social biases of the bert-base-uncased pre-training corpus (BookCorpus and English Wikipedia), which is known to encode gender, ethnic and occupational stereotypes. Nothing here mitigates or measures them.
  • Single seed. Every number on this card comes from one run with seed 42. No variance estimate exists, so a difference of roughly one point against another configuration must be read as noise rather than as an improvement.
  • Subsampling. Training used 14,041 examples (hyperparams.train_size). Read from a parquet mirror rather than from the legacy script-based conll2003, because datasets >= 4.0 removed dataset loading scripts.
  • Sequence truncation. Inputs were truncated to 128 tokens; longer inputs lose their tail at inference too.
  • Frozen-encoder caveat (feature-based models only). Not applicable: this model was fine-tuned, so the encoder weights moved.
  • Only a validation (and where available test) split was measured. No robustness, adversarial, out-of-domain or fairness evaluation was performed.

Environmental and compute footprint

Hardware 1 × Tesla T4 (Google Colab), CUDA 12.8
Training wall time 160.4 s (~2.7 min) over 1,317 steps
Seconds per step 0.122
Peak GPU memory 3076.6 MB
Software transformers 5.16.1 · torch 2.11.0+cu128 · datasets 4.8.5
Run timestamp 2026-09-25T06:44:42

A single short fine-tuning run on one T4 is a small footprint in absolute terms, and that is the point of this assignment: the cost-versus-quality table in the companion report shows how much of the quality survives when only a fraction of the parameters is trained, which is the part that scales. No CO₂e figure is claimed here, because the energy mix of the Colab region is unknown.

References and citation

  • Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL-HLT. arXiv:1810.04805. Section 5.3 is the feature-based versus fine-tuning comparison this model is part of.
  • Tjong Kim Sang, E. F., & De Meulder, F. (2003). Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition. CoNLL. https://aclanthology.org/W03-0419/
  • Tunstall, L., von Werra, L., & Wolf, T. (2022). Natural Language Processing with Transformers, chapters 1–3. O'Reilly.
  • Hugging Face, Fine-tuning a pretrained model: https://huggingface.co/docs/transformers/training
@inproceedings{devlin2019bert,
  title     = {BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding},
  author    = {Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina},
  booktitle = {Proceedings of NAACL-HLT},
  pages     = {4171--4186},
  year      = {2019},
  url       = {https://arxiv.org/abs/1810.04805}
}

Cite this adapted model as:

@misc{u2t01_ner_full,
  title  = {BERT adapted to named entity recognition (CoNLL-2003, IOB2) — full fine-tuning},
  author = {U2T01 team},
  year   = {2026},
  note   = {U2T01: Adapting BERT for NLP tasks. Adaptation method: full. Seed 42.},
  url    = {https://huggingface.co/abelcetina/u2t01-bert-conll2003-ner}
}
Downloads last month
22
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for abelcetina/u2t01-bert-conll2003-ner

Finetuned
(6999)
this model

Dataset used to train abelcetina/u2t01-bert-conll2003-ner

Paper for abelcetina/u2t01-bert-conll2003-ner