Mountain Named Entity Recognition (NER)

A BERT-based Named Entity Recognition model for identifying mountain names in text.

Overview

This project implements a custom NER solution to detect and extract mountain names from natural language text. The model is fine-tuned on a hybrid dataset combining real Wikipedia sentences and synthetically generated examples, achieving robust performance across various text styles and mountain name formats.

Key Features

  • βœ… BERT-based architecture (bert-base-cased)
  • βœ… Hybrid training dataset (80% Wikipedia, 20% synthetic)
  • βœ… BIO tagging scheme for sequence labeling
  • βœ… Multiple name format support (Mount Everest, Mt. Fuji, K2, etc.)
  • βœ… Ready-to-use inference script and Python API
  • βœ… Comprehensive documentation with Jupyter notebooks

Project Structure

Task 1. NER/
β”œβ”€β”€ data/                           # Dataset files
β”‚   β”œβ”€β”€ conll/                      # CoNLL format splits
β”‚   β”‚   β”œβ”€β”€ train.conll            # Training data (80%)
β”‚   β”‚   β”œβ”€β”€ val.conll              # Validation data (10%)
β”‚   β”‚   β”œβ”€β”€ test.conll             # Test data (10%)
β”‚   β”‚   └── mixed_sentences.conll  # Combined dataset
β”‚   β”œβ”€β”€ mountain_list.txt          # Curated mountain names
β”‚   β”œβ”€β”€ mountains_wiki_sentences.csv   # Wikipedia scraped data
β”‚   └── synthetic_sentences.csv    # Generated synthetic data
β”œβ”€β”€ models/                         # Trained model weights
β”‚   β”œβ”€β”€ config.json
β”‚   β”œβ”€β”€ model.safetensors
β”‚   β”œβ”€β”€ tokenizer.json
β”‚   └── ...
β”œβ”€β”€ notebooks/                      # Jupyter notebooks
β”‚   β”œβ”€β”€ dataset_creation.ipynb     # Dataset creation walkthrough
β”‚   └── demo.ipynb                 # Inference demo
β”œβ”€β”€ src/                            # Source code
β”‚   β”œβ”€β”€ mountains_wiki_scraper.py      # Wikipedia data scraper
β”‚   β”œβ”€β”€ mountains_dataset_generator.py # Synthetic data generator
β”‚   β”œβ”€β”€ prepare_conll_splits.py        # Dataset mixing & splitting
β”‚   β”œβ”€β”€ train.py                       # Model training script
β”‚   └── inference.py                   # Inference script with CLI
β”œβ”€β”€ requirements.txt                # Python dependencies
β”œβ”€β”€ potential_improvements.pdf      # Performance analysis & improvements
└── README.md                       # This file

Installation

Prerequisites

  • Python 3.8+
  • 8GB+ RAM recommended for training
  • CUDA-compatible GPU (optional, CPU supported)

Setup

  1. Clone or download the repository

  2. Create a virtual environment (recommended):

python -m venv venv
source venv/bin/activate  # On Windows: venv\Scripts\activate
  1. Install dependencies:
pip install -r requirements.txt

Quick Start

Using the Trained Model

Command Line Interface:

# Inference on a single sentence
python src/inference.py "Mount Everest is the highest peak in the world."

# Inference from file
python src/inference.py --file input.txt

# Simple output (mountains only)
python src/inference.py "Climbers visited K2 and Denali." --simple

Python API:

from src.inference import MountainNER

# Initialize model (loads from local or HuggingFace Hub)
ner = MountainNER()

# Extract mountain names
mountains = ner.extract_mountains("Mount Fuji is beautiful.")
print(mountains)  # ['Mount Fuji']

# Detailed predictions with BIO labels
ner.print_predictions("Climbers scaled Mount Rainier yesterday.")

Or use directly from HuggingFace:

from transformers import pipeline
ner = pipeline("ner", model="RomanTerendiy/RomanTerendiy")
print(ner("Mount Everest is the highest mountain."))

Interactive Demo

Explore the model with the demo notebook:

jupyter notebook notebooks/demo.ipynb

Dataset Creation

The dataset combines two complementary sources:

1. Wikipedia Scraping (~4000 sentences)

  • Real-world sentences from Wikipedia articles
  • Natural language diversity
  • Authentic context and usage patterns

2. Synthetic Generation (~1000 sentences)

  • Template-based sentence construction
  • Controlled linguistic patterns
  • Edge case coverage (Mt. vs Mount, etc.)

Data Pipeline

# Step 1: Scrape Wikipedia (optional, data already provided)
python src/mountains_wiki_scraper.py

# Step 2: Generate synthetic data (optional, data already provided)
python src/mountains_dataset_generator.py

# Step 3: Mix and split dataset
python src/prepare_conll_splits.py

See notebooks/dataset_creation.ipynb for a detailed walkthrough.

Model Training

Train from Scratch

python src/train.py

Training Configuration

Key hyperparameters in src/train.py:

  • Base model: bert-base-cased
  • Batch size: 2 (per device)
  • Gradient accumulation: 4 steps
  • Learning rate: 5e-5
  • Epochs: 4
  • Optimizer: AdamW with weight decay 0.01

Training Hardware

  • CPU: Supported (training takes ~1-2 hours)
  • GPU: Recommended for faster training (~15-20 minutes)

Model checkpoints are saved every epoch in models/checkpoint-*/.

Model Performance

The trained model achieves strong performance on the test set:

  • Precision: High accuracy in identified entities
  • Recall: Effective detection of mountain names
  • F1-Score: Balanced performance

Usage Examples

Example 1: Simple Detection

text = "Mount Everest is the highest mountain."
mountains = ner.extract_mountains(text)
# Output: ['Mount Everest']

Example 2: Multiple Mountains

text = "The team climbed K2 and then Mount Kilimanjaro."
mountains = ner.extract_mountains(text)
# Output: ['K2', 'Mount Kilimanjaro']

Example 3: Various Formats

text = "Mt. Fuji and Matterhorn are iconic peaks."
mountains = ner.extract_mountains(text)
# Output: ['Mt. Fuji', 'Matterhorn']

BIO Labeling Scheme

The model uses the standard BIO (Begin-Inside-Outside) tagging:

Label Description
B-MOUNTAIN Beginning of a mountain entity
I-MOUNTAIN Inside (continuation) of a mountain entity
O Outside any entity (regular word)

Example:

Sentence:  Mount  Everest  is  the  highest  peak  .
Labels:    B-MTN  I-MTN    O   O    O       O     O

Model Weights

Model Weights uploaded on HuggingFace Hub

  • Link: https://huggingface.co/RomanTerendiy/RomanTerendiy
  • Model ID: RomanTerendiy/RomanTerendiy
  • Direct loading:
    from transformers import AutoModelForTokenClassification, AutoTokenizer
    model = AutoModelForTokenClassification.from_pretrained("RomanTerendiy/RomanTerendiy")
    tokenizer = AutoTokenizer.from_pretrained("RomanTerendiy/RomanTerendiy")
    

Local Model Weights

  • Located in models/ directory
  • Automatically loaded by inference script if available

Using the Model Without Local Files

If you don't have the model weights locally, you can download directly from HuggingFace:

from src.inference import MountainNER

# Will automatically download from HuggingFace Hub
ner = MountainNER()  # Falls back to RomanTerendiy/RomanTerendiy if local not found
mountains = ner.extract_mountains("Mount Everest is the highest peak.")

Or load directly:

from transformers import pipeline
ner = pipeline("ner", model="RomanTerendiy/RomanTerendiy")

Potential Improvements

See potential_improvements.pdf for detailed discussion. Key areas:

Downloads last month
3
Safetensors
Model size
0.1B params
Tensor type
F32
Β·
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support