Mountain Named Entity Recognition (NER)
A BERT-based Named Entity Recognition model for identifying mountain names in text.
Overview
This project implements a custom NER solution to detect and extract mountain names from natural language text. The model is fine-tuned on a hybrid dataset combining real Wikipedia sentences and synthetically generated examples, achieving robust performance across various text styles and mountain name formats.
Key Features
- β BERT-based architecture (bert-base-cased)
- β Hybrid training dataset (80% Wikipedia, 20% synthetic)
- β BIO tagging scheme for sequence labeling
- β Multiple name format support (Mount Everest, Mt. Fuji, K2, etc.)
- β Ready-to-use inference script and Python API
- β Comprehensive documentation with Jupyter notebooks
Project Structure
Task 1. NER/
βββ data/ # Dataset files
β βββ conll/ # CoNLL format splits
β β βββ train.conll # Training data (80%)
β β βββ val.conll # Validation data (10%)
β β βββ test.conll # Test data (10%)
β β βββ mixed_sentences.conll # Combined dataset
β βββ mountain_list.txt # Curated mountain names
β βββ mountains_wiki_sentences.csv # Wikipedia scraped data
β βββ synthetic_sentences.csv # Generated synthetic data
βββ models/ # Trained model weights
β βββ config.json
β βββ model.safetensors
β βββ tokenizer.json
β βββ ...
βββ notebooks/ # Jupyter notebooks
β βββ dataset_creation.ipynb # Dataset creation walkthrough
β βββ demo.ipynb # Inference demo
βββ src/ # Source code
β βββ mountains_wiki_scraper.py # Wikipedia data scraper
β βββ mountains_dataset_generator.py # Synthetic data generator
β βββ prepare_conll_splits.py # Dataset mixing & splitting
β βββ train.py # Model training script
β βββ inference.py # Inference script with CLI
βββ requirements.txt # Python dependencies
βββ potential_improvements.pdf # Performance analysis & improvements
βββ README.md # This file
Installation
Prerequisites
- Python 3.8+
- 8GB+ RAM recommended for training
- CUDA-compatible GPU (optional, CPU supported)
Setup
Clone or download the repository
Create a virtual environment (recommended):
python -m venv venv
source venv/bin/activate # On Windows: venv\Scripts\activate
- Install dependencies:
pip install -r requirements.txt
Quick Start
Using the Trained Model
Command Line Interface:
# Inference on a single sentence
python src/inference.py "Mount Everest is the highest peak in the world."
# Inference from file
python src/inference.py --file input.txt
# Simple output (mountains only)
python src/inference.py "Climbers visited K2 and Denali." --simple
Python API:
from src.inference import MountainNER
# Initialize model (loads from local or HuggingFace Hub)
ner = MountainNER()
# Extract mountain names
mountains = ner.extract_mountains("Mount Fuji is beautiful.")
print(mountains) # ['Mount Fuji']
# Detailed predictions with BIO labels
ner.print_predictions("Climbers scaled Mount Rainier yesterday.")
Or use directly from HuggingFace:
from transformers import pipeline
ner = pipeline("ner", model="RomanTerendiy/RomanTerendiy")
print(ner("Mount Everest is the highest mountain."))
Interactive Demo
Explore the model with the demo notebook:
jupyter notebook notebooks/demo.ipynb
Dataset Creation
The dataset combines two complementary sources:
1. Wikipedia Scraping (~4000 sentences)
- Real-world sentences from Wikipedia articles
- Natural language diversity
- Authentic context and usage patterns
2. Synthetic Generation (~1000 sentences)
- Template-based sentence construction
- Controlled linguistic patterns
- Edge case coverage (Mt. vs Mount, etc.)
Data Pipeline
# Step 1: Scrape Wikipedia (optional, data already provided)
python src/mountains_wiki_scraper.py
# Step 2: Generate synthetic data (optional, data already provided)
python src/mountains_dataset_generator.py
# Step 3: Mix and split dataset
python src/prepare_conll_splits.py
See notebooks/dataset_creation.ipynb for a detailed walkthrough.
Model Training
Train from Scratch
python src/train.py
Training Configuration
Key hyperparameters in src/train.py:
- Base model:
bert-base-cased - Batch size: 2 (per device)
- Gradient accumulation: 4 steps
- Learning rate: 5e-5
- Epochs: 4
- Optimizer: AdamW with weight decay 0.01
Training Hardware
- CPU: Supported (training takes ~1-2 hours)
- GPU: Recommended for faster training (~15-20 minutes)
Model checkpoints are saved every epoch in models/checkpoint-*/.
Model Performance
The trained model achieves strong performance on the test set:
- Precision: High accuracy in identified entities
- Recall: Effective detection of mountain names
- F1-Score: Balanced performance
Usage Examples
Example 1: Simple Detection
text = "Mount Everest is the highest mountain."
mountains = ner.extract_mountains(text)
# Output: ['Mount Everest']
Example 2: Multiple Mountains
text = "The team climbed K2 and then Mount Kilimanjaro."
mountains = ner.extract_mountains(text)
# Output: ['K2', 'Mount Kilimanjaro']
Example 3: Various Formats
text = "Mt. Fuji and Matterhorn are iconic peaks."
mountains = ner.extract_mountains(text)
# Output: ['Mt. Fuji', 'Matterhorn']
BIO Labeling Scheme
The model uses the standard BIO (Begin-Inside-Outside) tagging:
| Label | Description |
|---|---|
B-MOUNTAIN |
Beginning of a mountain entity |
I-MOUNTAIN |
Inside (continuation) of a mountain entity |
O |
Outside any entity (regular word) |
Example:
Sentence: Mount Everest is the highest peak .
Labels: B-MTN I-MTN O O O O O
Model Weights
Model Weights uploaded on HuggingFace Hub
- Link: https://huggingface.co/RomanTerendiy/RomanTerendiy
- Model ID:
RomanTerendiy/RomanTerendiy - Direct loading:
from transformers import AutoModelForTokenClassification, AutoTokenizer model = AutoModelForTokenClassification.from_pretrained("RomanTerendiy/RomanTerendiy") tokenizer = AutoTokenizer.from_pretrained("RomanTerendiy/RomanTerendiy")
Local Model Weights
- Located in
models/directory - Automatically loaded by inference script if available
Using the Model Without Local Files
If you don't have the model weights locally, you can download directly from HuggingFace:
from src.inference import MountainNER
# Will automatically download from HuggingFace Hub
ner = MountainNER() # Falls back to RomanTerendiy/RomanTerendiy if local not found
mountains = ner.extract_mountains("Mount Everest is the highest peak.")
Or load directly:
from transformers import pipeline
ner = pipeline("ner", model="RomanTerendiy/RomanTerendiy")
Potential Improvements
See potential_improvements.pdf for detailed discussion. Key areas:
- Downloads last month
- 3