bert-base-conll03-ner
Fine-tuning of bert-base-uncased for Named Entity Recognition (NER), delivered as
part of assignment U2T01 (Adapting BERT for NLP tasks — Trends in Data Science,
Unit 2, Universidad Politécnica de Yucatán).
Model description
bert-base-uncased with a token-level linear classification head over the last
hidden state of every token, predicting one of 9 BIO-scheme entity tags (person,
organization, location, miscellaneous, or none).
Training data
- Dataset: CoNLL-2003
(
lhoestq/conll2003), using its own official train/validation/test splits. - Subword alignment: since the dataset labels whole words but the tokenizer can
split a word into multiple subwords, the label is placed on the first subword of
each word only; all other positions (continuation subwords,
[CLS],[SEP], padding) get label-100, whichCrossEntropyLossignores.
Training procedure
Two adaptation methods were trained and compared; full fine-tuning is the delivered model.
| Hyperparameter | Value |
|---|---|
| Base model | bert-base-uncased (110M params) |
| Method | Full fine-tuning (BERT body + head, jointly) |
| Learning rate (head) | 1e-3 |
| Learning rate (BERT body) | 2e-5 |
| Epochs | 4 |
| Batch size | 32 (train) / 64 (eval) |
| Seed | 42 |
| Trainable parameters | 108,898,569 |
| Training time | 4.4 min (single T4 GPU) |
Compared alternative (not delivered): partial fine-tuning, freezing the entire BERT body except its last 2 encoder layers (14,182,665 trainable params, 1.9 min).
Evaluation results
Metric: seqeval F1 (evaluated at the entity level, e.g. "New York" counts as one
LOC entity rather than two tokens — important because the dominant "O" class would
make per-token accuracy misleadingly high).
| Method | Val F1 | Test F1 | Test Precision | Test Recall |
|---|---|---|---|---|
| Partial fine-tuning (last 2 layers + head) | 88.75% | 85.37% | 83.58% | 87.23% |
| Full fine-tuning (delivered) | 94.27% | 90.16% | 89.36% | 90.97% |
The 4.79-point F1 gap on test is above the ±1–3 point noise margin. Note the consistent drop from validation to test in both methods — a known property of the CoNLL-2003 test split being harder than its validation split, not a sign of overfitting.
Intended uses & limitations
- Intended use: named entity recognition (person, organization, location, misc) on English news-style text similar to CoNLL-2003, for coursework/research.
- Limitations: trained only on CoNLL-2003 (Reuters newswire from the 1990s); may
not generalize well to informal text, social media, or domains with different
entity distributions (e.g. biomedical, legal). Single seed, 4 epochs, no extensive
hyperparameter search. Inherits biases from the source corpus and from
bert-base-uncased's pretraining data.
References
- Devlin, J., Chang, M. W., Lee, K., & Toutanova, K. (2018). BERT: Pre-training of Deep Bidirectional Transformers for Language Understanding. arXiv:1810.04805. https://arxiv.org/abs/1810.04805
- Hugging Face. Fine-tune a pretrained model. https://huggingface.co/docs/transformers/training
- Dataset: lhoestq/conll2003
- Downloads last month
- -
Model tree for Dalila-Ku/bert-base-conll03-ner
Base model
google-bert/bert-base-uncased