Assignment 2 — Yuval Malka Classification, Regression, Clustering & Evaluation Global Terrorism Database (GTD)

Video Presentation: https://youtu.be/6yeC_aRLvTU Part 1: Dataset Overview

For this assignment, I selected the Global Terrorism Database (GTD) from Kaggle. The dataset contains approximately 181,000 terrorism incidents recorded worldwide from 1970 to 2017. It includes 135 features, combining both numeric variables (casualty counts, coordinates, temporal attributes) and categorical descriptors (country, region, attack type, weapon type, target type, etc.).

The main regression question explored in this project was: Can we predict the number of wounded (nwound) in a terrorism event based on event characteristics such as location, weapon type, and attack type?

sample data preview

image

Part 2: Exploratory Data Analysis (EDA)

The EDA process focused on understanding the distribution of casualties, identifying missing values, detecting outliers, and exploring differences between attack types, regions, and weapon categories.

Key steps included:

Handling missing values: casualty-related NaNs were replaced with zeros; categorical missing values were set to “Unknown”.

Outlier detection: extreme casualty values were retained because they represent real, important events.

Converting latitude and longitude to numeric and removing invalid geographic entries.

Producing descriptive statistics for all numerical fields.

Examining correlations between important variables.

Visualizing severity across regions, weapon types, attack types, and time.

boxplot of nwound

download

download

download

download

download

download

Summary of EDA: The dataset is highly skewed, with most events causing zero or few casualties. Severity varies significantly across regions and attack types, indicating meaningful patterns for modeling. Geographic clustering shows concentrated hotspots of high-severity events. These observations guided feature engineering decisions in later stages.

Additional Research Questions

Several analytical questions were explored using visualizations:

How did the number of terrorist attacks change over time?

download

Which attack types are most common?

download

Which regions suffer the most average casualties?

download

Which weapon types tend to cause the highest severity?

download

Do suicide attacks cause more casualties compared to non-suicide attacks?

download

Are high-severity events geographically clustered?

download

These visual insights further demonstrated distinct severity patterns across event categories.

Part 3: Baseline Regression Model (Explanation)

The baseline Linear Regression model required for Part 3 is fully implemented later in Part 5, after feature engineering and PCA. This ordering ensures that the baseline is trained on a clean and meaningful feature set, rather than on raw, unprocessed data.

To keep the assignment structure complete, the Part 3 baseline code is shown below but not executed, since running it before preprocessing would produce inconsistent results. The evaluated baseline model — along with metrics (MAE, MSE, RMSE, R²) and comparisons — appears in Part 5 and satisfies all requirements of this section.

Part 4: Feature Engineering

This stage introduced several new features and transformations aimed at improving model performance.

Engineered numeric features included:

total_casualties (nkill + nwound)

severity_index (log-transformed total casualties)

attack_age (years since attack)

is_successful (binary outcome)

suicide_weapon_interaction (interaction term)

Categorical features were one-hot encoded (attack type, target type, weapon type, country, region), producing over 255 dummy variables.

Numeric features were standardized using StandardScaler.

Dimensionality reduction was then applied using PCA to condense the dataset into 50 principal components, reducing noise and improving model efficiency.

Unsupervised Learning: KMeans clustering (k=5) was performed on numeric features, and the resulting cluster_label was added as an engineered feature. PCA visualization of clusters revealed meaningful separation, suggesting that the clusters captured underlying event behavior patterns.

download

Summary: The combination of engineered features, scaling, encoding, PCA, and clustering produced a rich and informative feature set to support both regression and classification tasks.

Part 5: Regression Models — Training and Evaluation

Three regression models were trained using the engineered dataset:

Linear Regression (SGDRegressor)

Random Forest Regressor

Gradient Boosting Regressor

Before training, PCA-reduced features (50 components) were sampled to 20,000 rows for efficiency. An 80/20 train-test split was used.

Regression Results:

model comparison table

image

Summary of Results:

The Linear Regression baseline performed poorly because it could not capture nonlinear event patterns.

Random Forest significantly improved error metrics.

Gradient Boosting achieved the overall best performance with the lowest RMSE and highest R² score.

Winning regression model: Gradient Boosting Regressor This model was exported and uploaded as winning_regression_model.pkl.

download

Part 6: Exporting the Winning Regression Model

The Gradient Boosting model was saved using joblib and uploaded to the HuggingFace repository. File: winning_regression_model.pkl

Part 7: Regression-to-Classification Conversion

To convert the regression target (nwound) into categories, quantile-based thresholds were applied:

Class 0: Low severity (bottom 33%)

Class 1: Medium severity (33–66%)

Class 2: High severity (top 33%)

Due to the heavy skew toward zero casualties, quantile boundaries were computed manually.

balance distribution:

download

The dataset is imbalanced, particularly for the medium-severity class. Therefore metrics like recall and macro-F1 are more meaningful than accuracy.

Part 8: Classification Models — Training and Evaluation

Using the PCA-reduced feature matrix, three different classification models were trained:

Logistic Regression

Random Forest Classifier

HistGradientBoostingClassifier

Each model was evaluated using classification reports and confusion matrices.

logistic regression confusion matrix random forest confusion matrix HistGradientBoosting confusion matrix

download

download

download

Summary of Findings:

Logistic Regression performed strongly on low- and high-severity classes but struggled with medium severity.

Random Forest produced more errors across classes and was less stable.

HistGradientBoostingClassifier performed the most consistently and accurately across all classes, especially reducing errors in the medium and high categories.

Winning classification model: HistGradientBoostingClassifier This model was exported as winning_classification_model.pkl.

Part 9: Presentation Video

The video presentation includes:

An introduction to the dataset and research question

Key EDA insights and visualizations

The full feature engineering process

Description of PCA and clustering

Regression model training and comparison

Classification model training and evaluation

Lessons learned and reflections

Part 10: Repository Contents

This repository includes:

README file

Jupyter Notebook (.ipynb)

winning_regression_model.pkl

winning_classification_model.pkl

Presentation video link

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support