Instructions to use abelcetina/u2t01-bert-conll2003-ner with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use abelcetina/u2t01-bert-conll2003-ner with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="abelcetina/u2t01-bert-conll2003-ner")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("abelcetina/u2t01-bert-conll2003-ner") model = AutoModelForTokenClassification.from_pretrained("abelcetina/u2t01-bert-conll2003-ner", device_map="auto") - Notebooks
- Google Colab
- Kaggle
BERT adapted to named entity recognition (CoNLL-2003, IOB2) — full fine-tuning
abelcetina/u2t01-bert-conll2003-ner — bert-base-uncased adapted to named entity recognition (CoNLL-2003, IOB2) with full fine-tuning.
Produced for the university assignment U2T01: Adapting BERT for NLP tasks. The companion repository contains the training code, the raw results file behind every number below, and the comparison against the alternative adaptation methods.
Model description
- Base encoder:
bert-base-uncased(108,898,569 parameters in total). - Task: named entity recognition (CoNLL-2003, IOB2).
- Adaptation method: full fine-tuning. Every parameter of the encoder and of the task head received gradients, with a small learning rate on the body and a larger one on the randomly initialised head.
- Head: a linear token-classification head over every token representation.
- Label set: 9 IOB2 tags: O, B-PER, I-PER, B-ORG, I-ORG, B-LOC, I-LOC, B-MISC, I-MISC
- Trained by: U2T01 team · generated 2026-09-25.
Intended use
- Research and teaching: reproducing the feature-based versus fine-tuning comparison of BERT section 5.3 across four task types.
- Inference on English text that resembles the training domain (Reuters newswire text from 1996–1997).
- Usage:
from transformers import pipeline
tagger = pipeline("token-classification", model="abelcetina/u2t01-bert-conll2003-ner", aggregation_strategy="simple")
print(tagger("Angela Merkel visited the European Central Bank in Frankfurt."))
Out-of-scope use
- Any decision affecting a person (hiring, moderation with consequences, credit, legal or medical decisions). This is a coursework model validated on one academic benchmark with a single seed.
- Languages other than English, and domains far from Reuters newswire text from 1996–1997 — performance is not characterised there and is expected to drop.
- Extracting facts about the world: for the QA model, the answer is a span copied from the context passage it is given, never knowledge the model holds.
- Commercial redistribution without checking the training-data licence below.
Training data
| Dataset | CoNLL-2003 (parquet mirror) |
| Identifier used in code | lhoestq/conll2003 |
| Licence | other — derived from the Reuters-21578 news corpus; research use under the Reuters terms referenced on the dataset card |
| Splits | train 14,041 · validation 3,250 · test 3,453 sentences |
| Training examples actually used | 14,041 |
| Evaluation split reported below | validation |
Read from a parquet mirror rather than from the legacy script-based conll2003, because datasets >= 4.0 removed dataset loading scripts.
Adaptation method and what was frozen
- Method flag:
--method full - Trainable components: the whole BERT encoder (embeddings + 12 layers) and the task head
- Frozen components: nothing
| Parameters | Value |
|---|---|
| Total | 108,898,569 |
| Trainable | 108,898,569 |
| Trainable share | 100 % |
Hyper-parameters
| Hyper-parameter | Value |
|---|---|
| Epochs | 3 |
| Batch size | 32 |
| Head learning rate | 0.001 |
| Body learning rate | 0.00002 |
| Weight decay | 0.01 |
| Max sequence length | 128 |
| Top layers unfrozen | all 12 encoder layers (full fine-tuning) |
| Seed | 42 |
Evaluation results
Metrics computed by src/metrics.py; the raw values live in ner__full__seed42.json.
Evaluation split (validation)
| Metric | Value |
|---|---|
loss |
0.0462 |
precision |
0.9338 |
recall |
0.9453 |
f1 |
0.9395 |
accuracy |
0.9882 |
Test split
| Metric | Value |
|---|---|
loss |
0.1009 |
precision |
0.8937 |
recall |
0.9102 |
f1 |
0.9019 |
accuracy |
0.9805 |
Headline number: f1 = 0.9395 (accuracy = 0.9882).
The comparison against the other adaptation methods of this task, with the gap in percentage points and the noise flag, is in report/results_tables.md of the companion code repository (Table 1, Table 2.ner and Table 3). A gap below one point there is reported as noise, not as an improvement.
Limitations and bias
- Dataset bias. Reuters newswire from the late 1990s: person, organisation and location names are heavily skewed towards the actors and the spelling conventions of that period and towards sports reporting, which dominates the corpus. The model inherits both that domain skew and the social biases
of the
bert-base-uncasedpre-training corpus (BookCorpus and English Wikipedia), which is known to encode gender, ethnic and occupational stereotypes. Nothing here mitigates or measures them. - Single seed. Every number on this card comes from one run with seed
42. No variance estimate exists, so a difference of roughly one point against another configuration must be read as noise rather than as an improvement. - Subsampling. Training used 14,041 examples (
hyperparams.train_size). Read from a parquet mirror rather than from the legacy script-basedconll2003, becausedatasets >= 4.0removed dataset loading scripts. - Sequence truncation. Inputs were truncated to 128 tokens; longer inputs lose their tail at inference too.
- Frozen-encoder caveat (feature-based models only). Not applicable: this model was fine-tuned, so the encoder weights moved.
- Only a validation (and where available test) split was measured. No robustness, adversarial, out-of-domain or fairness evaluation was performed.
Environmental and compute footprint
| Hardware | 1 × Tesla T4 (Google Colab), CUDA 12.8 |
| Training wall time | 160.4 s (~2.7 min) over 1,317 steps |
| Seconds per step | 0.122 |
| Peak GPU memory | 3076.6 MB |
| Software | transformers 5.16.1 · torch 2.11.0+cu128 · datasets 4.8.5 |
| Run timestamp | 2026-09-25T06:44:42 |
A single short fine-tuning run on one T4 is a small footprint in absolute terms, and that is the point of this assignment: the cost-versus-quality table in the companion report shows how much of the quality survives when only a fraction of the parameters is trained, which is the part that scales. No CO₂e figure is claimed here, because the energy mix of the Colab region is unknown.
References and citation
- Devlin, J., Chang, M.-W., Lee, K., & Toutanova, K. (2019). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. NAACL-HLT. arXiv:1810.04805. Section 5.3 is the feature-based versus fine-tuning comparison this model is part of.
- Tjong Kim Sang, E. F., & De Meulder, F. (2003). Introduction to the CoNLL-2003 Shared Task: Language-Independent Named Entity Recognition. CoNLL. https://aclanthology.org/W03-0419/
- Tunstall, L., von Werra, L., & Wolf, T. (2022). Natural Language Processing with Transformers, chapters 1–3. O'Reilly.
- Hugging Face, Fine-tuning a pretrained model: https://huggingface.co/docs/transformers/training
@inproceedings{devlin2019bert,
title = {BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding},
author = {Devlin, Jacob and Chang, Ming-Wei and Lee, Kenton and Toutanova, Kristina},
booktitle = {Proceedings of NAACL-HLT},
pages = {4171--4186},
year = {2019},
url = {https://arxiv.org/abs/1810.04805}
}
Cite this adapted model as:
@misc{u2t01_ner_full,
title = {BERT adapted to named entity recognition (CoNLL-2003, IOB2) — full fine-tuning},
author = {U2T01 team},
year = {2026},
note = {U2T01: Adapting BERT for NLP tasks. Adaptation method: full. Seed 42.},
url = {https://huggingface.co/abelcetina/u2t01-bert-conll2003-ner}
}
- Downloads last month
- 22
Model tree for abelcetina/u2t01-bert-conll2003-ner
Base model
google-bert/bert-base-uncased