PHMSA incident models

Two DistilBERT models fine-tuned on public PHMSA pipeline incident reports (gas distribution, gas transmission and gathering, hazardous liquid; 2010 to present). Used by the Space phmsa-incident-extraction.

Folder Task Held-out result
ner/ token classification, 12 entity types micro-F1 0.608 on 840 spans (1,050 sentences labelled by one annotator)
severity/ narrative classification: minor, moderate, severe, critical macro-F1 0.777, accuracy 0.956 on 1,921 narratives

Limits

  • Severity labels come from PHMSA's structured flags (fatality, injury requiring inpatient hospitalization, ignition, explosion), not from the narrative text. About 40% of critical narratives never mention a death (rough keyword estimate) and the model detects none of those. Severity says nothing about spill size or environmental damage.
  • The entity model does not extract root causes (CAUSE_FACTOR scores 0) and is weak on locations and regulatory references. Rare classes per commodity are too small to evaluate.
  • Tested only on PHMSA narratives. Not for operational, safety or regulatory decisions.

Loading

Download the repo with huggingface_hub.snapshot_download("Chinonso11/phmsa-incident-models"), then load the two subfolders with AutoModelForTokenClassification.from_pretrained(<path>/ner) and AutoModelForSequenceClassification.from_pretrained(<path>/severity).

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support

Model tree for Chinonso11/phmsa-incident-models

Finetuned
(12680)
this model