Model Card for Model ID

Model Details

Model Description

This entity recognition model is used for information extraction in job-related texts. It can identify entities of Occupation, Skills, Qualifications, Experience and Domain. It is trained based on the open-source dataset provided by Green et. al.

  • Developed by: International Hellenic University NLP
  • Model type: Token Classification
  • Language(s) (NLP): English
  • Finetuned from model: FacebookAI/roberta-base

Uses

Load model directly

from transformers import AutoTokenizer, AutoModelForTokenClassification

tokenizer = AutoTokenizer.from_pretrained("tabiya/roberta-base-job-ner")
model = AutoModelForTokenClassification.from_pretrained("tabiya/roberta-base-job-ner")

Training Details

Training Data

More information about the training dataset can be found here

Training Procedure

The training of this model was done using the HuggingFace token classification tutorial

Training Hyperparameters

  • Max Length: 128
  • Batch Size: 32
  • Learning Rate: 0.0001
  • Epochs: 5
  • Weight Decay: 0.01

Evaluation

Testing Data, Factors & Metrics

Testing Data

The training was evaluated on the test set of the training dataset provided by Green et. al.

Metrics

Following standard procedures for evaluating entity recognition training the chosen metric was the strict span-F1 measure provided by the seqeval library.

Results

Entity Category Strict Span-F1
Domain 0.3
Experience 0.56
Occuption 0.83
Qualification 0.55
Skill 0.51
Micro Average 0.56

Compute Infrastructure

The hyperparameter search and training was done on the Advanced Research Computing (ARC) of the University of Oxford.

Hardware

All training was performed on V100 GPUs.

Citation

BibTeX:

TBD

Downloads last month
134
Safetensors
Model size
0.1B params
Tensor type
F32
·
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support