Instructions to use Mudi12137/WaterBERT-NER with libraries, inference providers, notebooks, and local apps. Follow these links to get started.
- Libraries
- Transformers
How to use Mudi12137/WaterBERT-NER with Transformers:
# Use a pipeline as a high-level helper from transformers import pipeline pipe = pipeline("token-classification", model="Mudi12137/WaterBERT-NER")# Load model directly from transformers import AutoTokenizer, AutoModelForTokenClassification tokenizer = AutoTokenizer.from_pretrained("Mudi12137/WaterBERT-NER") model = AutoModelForTokenClassification.from_pretrained("Mudi12137/WaterBERT-NER", device_map="auto") - Notebooks
- Google Colab
- Kaggle
WaterBERT-NER
Named-entity recognition for wastewater- and water-treatment literature, fine-tuned from WaterBERT. It tags six entity types with a BIO scheme (13 labels):
| Entity type | Covers (examples) |
|---|---|
Pollutant |
contaminants and water-quality targets โ ammonium, sulfamethoxazole, COD, microplastics |
Wastewater_Treatment_Process |
treatment processes and unit operations โ adsorption, Fenton process, anaerobic digestion |
Reactor |
reactors, units and devices โ sequencing batch reactor, membrane bioreactor, constructed wetland |
Treatment_Parameter |
operating, reaction and material parameters โ pH, hydraulic retention time, specific surface area |
Microorganism |
microorganisms and microbial groups โ Nitrosomonas, anammox bacteria, microalgae |
Dosed_Material |
added materials, catalysts, adsorbents and reagents โ biochar, persulfate, PAC |
This is the model used to build the WaterKG entity graph.
Usage
from transformers import pipeline
ner = pipeline("token-classification", model="Mudi12137/WaterBERT-NER", aggregation_strategy="simple")
ner("A sequencing batch reactor inoculated with Nitrosomonas removed 95% of ammonium and "
"sulfamethoxazole at pH 7.5 after dosing biochar.")
# Reactor: sequencing batch reactor | Microorganism: nitrosomonas | Pollutant: ammonium,
# sulfamethoxazole | Treatment_Parameter: ph | Dosed_Material: biochar
Inputs longer than 512 tokens are truncated; split long documents into sentences or passages.
The tokenizer lower-cases input, so returned word strings are lower-case โ use the
start/end character offsets to recover the original spelling.
Labels
O, then B-/I- for Pollutant, Wastewater_Treatment_Process, Reactor,
Treatment_Parameter, Microorganism, Dosed_Material (see label_map.json).
Training
Fine-tuned from WaterBERT on manually annotated water-treatment abstracts; the annotation guideline and data are described in the accompanying paper.
Limitations
- Numeric values and removal efficiencies are not entity types of this model. The relation model WaterBERT-RE expects value spans from a separate value detector.
- Boundaries of long compound names (e.g. "rotating disc electrocoagulation system") and the type of cross-category words (e.g. adsorption as a process vs. a parameter) are the most common sources of error.
License
Apache-2.0.
- Downloads last month
- 11