LLMCPNER

Model checkpoint for the manuscript "LLMCPNER: Integrating Large Language Models and Curriculum Learning for Paleontological Named Entity Recognition", currently under review.

LLMCPNER is a span-based named entity recognition model for paleontological literature. It combines SciBERT, multi-model voting, and confidence-weighted curriculum learning.

Entity types

The model recognizes seven entity types: taxa, location, section, strata, lithology, facies, and age.

Evaluation

The model was evaluated on a manually corrected test set containing 188 texts and 1,348 entities.

Matching criterion Precision Recall F1
Strict 88.19 87.54 87.86
Partial 91.26 90.58 90.92

Strict matching requires exact entity boundaries and type. Partial matching requires at least 50% boundary overlap and the correct entity type.

Files

  • model.pt: PyTorch state dictionary of the final model.
  • label_mapping.json: entity label-to-ID and ID-to-label mappings.
  • training_config.json: core settings recorded for the reported run.

This checkpoint uses a custom span-classification architecture and is not a drop-in AutoModel.from_pretrained() model. The implementation and test set are available in the associated code repository:

https://github.com/goodXHD/LLMCPNER-Integrating-Large-Language-Models-and-Curriculum-Learning-for-Paleontological-Named-Entity

Intended use

The model is intended for research on named entity recognition in English paleontological literature. Performance outside this domain has not been established. Predictions should be reviewed before use in scientific databases or downstream knowledge resources.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for xhd521/LLMCPNER

Finetuned
(128)
this model