YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
- Water Systems Attack Detection β CAIA 2022
- π Project Overview
- Accuracy
- Precision
- Recall
- F1-Score
- Confusion Matrix
- Machine Learning
- Cybersecurity
- Critical Infrastructure
- 1. Load the Dataset
- 2. Explore the Data
- 3. Prepare Features
- 4. Prepare Labels
- 5. Train the Model
- 6. Test the Model
- 7. Evaluate
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- π Project Overview
- π― Project Objective
- Accuracy
- Precision
- Recall
- F1-Score
- Confusion Matrix
- Machine Learning
- Cybersecurity
- Critical Infrastructure
- 1. Load the Dataset
- 2. Explore the Data
- 3. Prepare Features
- 4. Prepare Labels
- 5. Train the Model
- 6. Test the Model
- 7. Evaluate
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- Accuracy
- π‘οΈ Why Water-System Security?
- Accuracy
- Precision
- Recall
- F1-Score
- Confusion Matrix
- Machine Learning
- Cybersecurity
- Critical Infrastructure
- 1. Load the Dataset
- 2. Explore the Data
- 3. Prepare Features
- 4. Prepare Labels
- 5. Train the Model
- 6. Test the Model
- 7. Evaluate
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- Accuracy
- π¬ Problem Definition
- Accuracy
- Precision
- Recall
- F1-Score
- Confusion Matrix
- Machine Learning
- Cybersecurity
- Critical Infrastructure
- 1. Load the Dataset
- 2. Explore the Data
- 3. Prepare Features
- 4. Prepare Labels
- 5. Train the Model
- 6. Test the Model
- 7. Evaluate
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- Accuracy
- π Dataset
- Accuracy
- Precision
- Recall
- F1-Score
- Confusion Matrix
- Machine Learning
- Cybersecurity
- Critical Infrastructure
- 1. Load the Dataset
- 2. Explore the Data
- 3. Prepare Features
- 4. Prepare Labels
- 5. Train the Model
- 6. Test the Model
- 7. Evaluate
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- Accuracy
- π§Ή Data Preprocessing
- Accuracy
- Precision
- Recall
- F1-Score
- Confusion Matrix
- Machine Learning
- Cybersecurity
- Critical Infrastructure
- 1. Load the Dataset
- 2. Explore the Data
- 3. Prepare Features
- 4. Prepare Labels
- 5. Train the Model
- 6. Test the Model
- 7. Evaluate
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- Accuracy
- π§ Machine Learning Approach
- Accuracy
- Precision
- Recall
- F1-Score
- Confusion Matrix
- Machine Learning
- Cybersecurity
- Critical Infrastructure
- 1. Load the Dataset
- 2. Explore the Data
- 3. Prepare Features
- 4. Prepare Labels
- 5. Train the Model
- 6. Test the Model
- 7. Evaluate
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- Accuracy
- βοΈ Model Pipeline
- Accuracy
- Precision
- Recall
- F1-Score
- Confusion Matrix
- Machine Learning
- Cybersecurity
- Critical Infrastructure
- 1. Load the Dataset
- 2. Explore the Data
- 3. Prepare Features
- 4. Prepare Labels
- 5. Train the Model
- 6. Test the Model
- 7. Evaluate
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- Accuracy
- π¨ Attack Detection
- Accuracy
- Precision
- Recall
- F1-Score
- Confusion Matrix
- Machine Learning
- Cybersecurity
- Critical Infrastructure
- 1. Load the Dataset
- 2. Explore the Data
- 3. Prepare Features
- 4. Prepare Labels
- 5. Train the Model
- 6. Test the Model
- 7. Evaluate
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- Accuracy
- π Model Evaluation
- Accuracy
- Precision
- Recall
- F1-Score
- Confusion Matrix
- Machine Learning
- Cybersecurity
- Critical Infrastructure
- 1. Load the Dataset
- 2. Explore the Data
- 3. Prepare Features
- 4. Prepare Labels
- 5. Train the Model
- 6. Test the Model
- 7. Evaluate
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- Accuracy
- π Important Security Concepts
- Machine Learning
- Cybersecurity
- Critical Infrastructure
- 1. Load the Dataset
- 2. Explore the Data
- 3. Prepare Features
- 4. Prepare Labels
- 5. Train the Model
- 6. Test the Model
- 7. Evaluate
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- Machine Learning
- π Industrial Control System Context
- 1. Load the Dataset
- 2. Explore the Data
- 3. Prepare Features
- 4. Prepare Labels
- 5. Train the Model
- 6. Test the Model
- 7. Evaluate
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- 1. Load the Dataset
- π§ͺ Experimental Workflow
- 1. Load the Dataset
- 2. Explore the Data
- 3. Prepare Features
- 4. Prepare Labels
- 5. Train the Model
- 6. Test the Model
- 7. Evaluate
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- 1. Load the Dataset
- π Class Distribution
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- Programming
- π§ Why Machine Learning?
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- Programming
- π¨βπ» My Contribution
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- Programming
- π οΈ Technologies
- Programming
- Data Processing
- Machine Learning
- Visualization
- Development Environment
- 1. Open the Notebook
- 2. Load the Dataset
- 3. Run Preprocessing
- 4. Train the Model
- 5. Evaluate
- 6. Analyze Results
- Dataset Dependency
- Class Imbalance
- False Positives
- False Negatives
- Dataset-to-Real-World Gap
- π¨βπ» Modified by Momen
- Programming
- π¦ Installation
- π Running the Project
- π Suggested Repository Structure
- π¬ Research Perspective
- π Real-World Applications
- β οΈ Limitations
- π Future Improvements
- π Learning Outcomes
- π Keywords
- π Disclaimer
- π Project Summary
Water Systems Attack Detection β CAIA 2022
π Project Overview
Water Systems Attack Detection β CAIA 2022 is a machine-learning project focused on detecting cyber-attacks and abnormal behavior in water distribution systems.
The project is based on work related to the Virginia Tech / CAIA 2022 research context, where machine-learning techniques are applied to identify malicious or abnormal activity within water-system operational data.
The main goal is to build a model capable of distinguishing between:
- π’ Normal system behavior
- π΄ Attack / anomalous behavior
This type of system is part of Industrial Control System (ICS) security and Cyber-Physical System (CPS) security, where attacks against digital control systems can potentially affect physical infrastructure.
π― Project Objective
The primary objective of this project is:
Develop a machine-learning-based intrusion detection approach capable of identifying attacks targeting water distribution system operations.
The project aims to demonstrate how data-driven models can analyze sensor and system measurements and identify patterns that differ from normal operational behavior.
The overall concept can be summarized as:
Water System Data
β
βΌ
Data Preprocessing
β
βΌ
Feature Engineering
β
βΌ
Machine Learning Model
β
βΌ
Attack Detection
β
βββββ΄βββββ
βΌ βΌ
Normal Attack
π‘οΈ Why Water-System Security?
Modern water infrastructure increasingly relies on:
- Sensors
- Programmable Logic Controllers (PLCs)
- Supervisory Control and Data Acquisition (SCADA)
- Industrial communication networks
- Automated control systems
- Digital monitoring
This creates a connection between the physical water infrastructure and the cyber environment.
Consequently, an attacker who compromises a control or monitoring system could potentially manipulate system measurements or operational commands.
Machine learning can provide an additional security layer by learning the normal behavior of the system and detecting deviations.
π¬ Problem Definition
Traditional security mechanisms often rely on predefined signatures or rules.
However, industrial environments can experience:
Unknown Attacks
+
Changing System Behavior
+
Large Sensor Data
β
Difficulty Detecting Anomalies
A machine-learning-based detector can instead learn patterns from historical system data.
The model attempts to learn:
Normal Behavior
β
Expected Patterns
β
Compare New Observation
β
Deviation Detected?
β
ββββ΄βββ
β β
NO YES
β β
Normal Attack
π Dataset
The project works with data representing the behavior of a water system under normal and attack conditions.
The data can contain measurements generated from different components of the water infrastructure, such as:
- Water-system sensors
- Process measurements
- System states
- Operational variables
- Control-related measurements
- Attack indicators / labels
The exact feature configuration depends on the specific dataset version used with the notebook.
π§Ή Data Preprocessing
Before training the machine-learning model, the dataset is prepared for learning.
A typical preprocessing pipeline includes:
Raw Dataset
β
βΌ
Data Loading
β
βΌ
Missing / Invalid Values
β
βΌ
Feature Selection
β
βΌ
Data Cleaning
β
βΌ
Feature Scaling
β
βΌ
Train / Test Split
β
βΌ
Machine Learning
Numerical features may require normalization or standardization because different sensors can operate on significantly different scales.
π§ Machine Learning Approach
The model treats attack detection as a supervised classification problem when attack labels are available.
The general formulation is:
X = System Measurements
y = System State / Attack Label
The classifier learns:
X β Model β y
where:
y = Normal
or:
y = Attack
Depending on the specific experiment, the classification problem may also contain multiple attack classes.
βοΈ Model Pipeline
The complete workflow can be represented as:
Water-System Dataset
β
βΌ
Data Preprocessing
β
βΌ
Feature Selection
β
βΌ
Normalization
β
βΌ
Train / Test
β
ββββββββββββ΄βββββββββββ
β β
βΌ βΌ
Training Testing
β β
βΌ βΌ
Machine Learning Model Predictions
β β
ββββββββββββ¬βββββββββββ
βΌ
Evaluation
β
ββββββββββββ΄βββββββββββ
βΌ βΌ
Normal Attack
π¨ Attack Detection
The central task is identifying whether a particular observation represents normal system behavior or an attack.
Conceptually:
New Water-System Observation
β
βΌ
Feature Vector
β
βΌ
Trained Classifier
β
βΌ
Attack Probability
β
βββββββ΄ββββββ
βΌ βΌ
Normal Attack
This approach can help security systems detect suspicious behavior before it becomes a larger operational problem.
π Model Evaluation
The model should not be evaluated using accuracy alone.
For cybersecurity and attack detection, several metrics are important:
Accuracy
Measures the percentage of correctly classified samples.
Accuracy =
Correct Predictions / Total Predictions
Precision
Measures how many samples predicted as attacks were actually attacks.
Precision =
TP / (TP + FP)
Recall
Measures how many actual attacks were successfully detected.
Recall =
TP / (TP + FN)
F1-Score
Provides a balance between precision and recall.
F1 =
2 Γ Precision Γ Recall
-----------------------
Precision + Recall
Confusion Matrix
A confusion matrix provides a detailed view of classification performance:
Predicted
Normal Attack
ββββββββββ¬βββββββββ
Actual Normal β TN β FP β
ββββββββββΌβββββββββ€
Actual Attack β FN β TP β
ββββββββββ΄βββββββββ
For an attack-detection system, false negatives are particularly important, because an undetected attack can represent a security risk.
π Important Security Concepts
This project demonstrates several concepts at the intersection of:
Machine Learning
- Supervised learning
- Classification
- Feature engineering
- Data preprocessing
- Model evaluation
Cybersecurity
- Intrusion detection
- Anomaly detection
- Attack classification
- Industrial cybersecurity
- Cyber-Physical Systems
Critical Infrastructure
- Water-system monitoring
- Industrial control systems
- Sensor-based systems
- SCADA environments
π Industrial Control System Context
The project can be viewed as an ICS intrusion-detection problem.
A simplified architecture is:
Physical Water System
β
βΌ
Sensors
β
βΌ
PLCs
β
βΌ
SCADA / Control Layer
β
βΌ
Network Traffic
β
βΌ
Security Monitoring
β
βΌ
Machine Learning Model
β
βββββββ΄ββββββ
βΌ βΌ
Normal Attack
The machine-learning component therefore acts as an additional monitoring layer around the industrial system.
π§ͺ Experimental Workflow
The project follows a typical machine-learning experimentation process:
1. Load the Dataset
Dataset
β
DataFrame / NumPy
2. Explore the Data
Feature Distribution
+
Class Distribution
+
Missing Values
β
Data Understanding
3. Prepare Features
Raw Features
β
Cleaning
β
Scaling
β
Feature Matrix X
4. Prepare Labels
Attack / Normal
β
Target y
5. Train the Model
X_train + y_train
β
Machine Learning
β
Trained Model
6. Test the Model
X_test
β
Model
β
Predictions
7. Evaluate
Predictions
β
Accuracy
Precision
Recall
F1-Score
Confusion Matrix
π Class Distribution
Before training, class distribution should be inspected.
For example:
Dataset
β
βββ Normal Samples
β
βββ Attack Samples
If the classes are highly imbalanced, accuracy may provide a misleading evaluation.
Therefore, metrics such as precision, recall, F1-score, and confusion matrix become especially important.
π§ Why Machine Learning?
Machine learning is useful in this context because system behavior can contain complex patterns that are difficult to represent using manually written rules.
Instead of explicitly defining:
IF sensor_A > X
AND sensor_B < Y
AND sensor_C = Z
THEN attack
the model can learn relationships from historical examples:
Historical System Data
β
βΌ
ML Training
β
βΌ
Learned System Patterns
β
βΌ
New System Observation
β
βΌ
Attack Detection
π¨βπ» My Contribution
This project was reviewed, modified, and adapted by Momen.
I worked on the implementation and experimentation to make the original approach more structured and suitable for practical machine-learning experimentation.
My modifications focus on areas such as:
- Organizing the notebook workflow
- Improving preprocessing
- Structuring the machine-learning pipeline
- Running model experiments
- Evaluating predictions
- Improving visualization and analysis
- Making the implementation easier to understand and reproduce
The project should therefore be considered an adapted and modified implementation, rather than a claim that the underlying research or dataset was originally created by me.
π οΈ Technologies
The project is primarily implemented using Python and common machine-learning tools.
Programming
- Python
Data Processing
- NumPy
- Pandas
Machine Learning
- Scikit-learn
- Machine-learning classification algorithms
Visualization
- Matplotlib
- Seaborn
Development Environment
- Jupyter Notebook
- Google Colab
π¦ Installation
Clone the repository or download the notebook and install the required dependencies.
pip install numpy pandas matplotlib seaborn scikit-learn jupyter
Depending on the exact notebook implementation, additional libraries may be required.
π Running the Project
1. Open the Notebook
jupyter notebook
or use Google Colab.
2. Load the Dataset
Place the water-system dataset in the expected dataset directory.
3. Run Preprocessing
Execute the preprocessing cells to:
- Load the data
- Clean the dataset
- Select features
- Prepare labels
- Split the data
4. Train the Model
Run the training cells to generate the trained classifier.
5. Evaluate
Evaluate the model using:
Accuracy
Precision
Recall
F1-Score
Confusion Matrix
6. Analyze Results
Inspect:
Predictions
+
Classification Metrics
+
Confusion Matrix
+
Visualizations
π Suggested Repository Structure
Water_Systems_Attack_Detection_CAIA_2022_Virginia_Tech/
β
βββ README.md
β
βββ notebooks/
β βββ water_system_attack_detection.ipynb
β
βββ data/
β βββ README.md
β
βββ models/
β βββ trained_models/
β
βββ results/
β βββ figures/
β βββ predictions/
β βββ metrics/
β
βββ requirements.txt
β
βββ .gitignore
π¬ Research Perspective
The project demonstrates an important research direction:
Cybersecurity
+
Machine Learning
+
Critical Infrastructure
β
Intelligent Attack Detection
Rather than treating cybersecurity and machine learning as separate disciplines, the project demonstrates how ML can be integrated into security monitoring for cyber-physical infrastructure.
π Real-World Applications
The same general methodology can potentially be applied to other critical infrastructure environments, including:
- Water treatment systems
- Water distribution networks
- Power grids
- Manufacturing systems
- Industrial plants
- Oil and gas infrastructure
- Smart infrastructure
- SCADA environments
The model itself, however, should be validated separately for each environment because system behavior and attack characteristics can differ significantly.
β οΈ Limitations
Machine-learning-based attack detection has several limitations.
Dataset Dependency
A model can learn patterns specific to its training dataset and may not generalize perfectly to another water system.
Class Imbalance
Attack datasets may contain significantly fewer attack samples than normal samples.
False Positives
Normal operational changes can sometimes be incorrectly classified as attacks.
False Negatives
Some attacks may resemble legitimate system behavior and therefore remain undetected.
Dataset-to-Real-World Gap
Performance on a benchmark dataset does not automatically guarantee performance on a real operational water infrastructure environment.
π Future Improvements
Potential improvements include:
- Compare multiple machine-learning algorithms
- Add Random Forest and Gradient Boosting
- Experiment with XGBoost
- Test neural-network-based classifiers
- Add anomaly-detection models
- Perform feature selection
- Address class imbalance
- Add cross-validation
- Tune model hyperparameters
- Add ROC-AUC and Precision-Recall curves
- Analyze false positives and false negatives
- Evaluate robustness against previously unseen attacks
- Investigate explainable AI techniques
- Compare supervised and unsupervised approaches
- Test temporal/deep-learning models such as LSTM
- Develop a real-time attack detection pipeline
π Learning Outcomes
Through this project, the following concepts are explored:
- Machine-learning classification
- Cybersecurity analytics
- Industrial Control System security
- Cyber-Physical Systems
- Water infrastructure security
- Data preprocessing
- Feature engineering
- Classification metrics
- Confusion matrices
- Attack detection
- Anomaly detection
- Critical infrastructure protection
- Machine-learning-based intrusion detection
π Keywords
Water Systems
Water Distribution System
Attack Detection
Cybersecurity
Industrial Control Systems
ICS Security
SCADA
Cyber-Physical Systems
Machine Learning
Intrusion Detection
Anomaly Detection
Critical Infrastructure
Virginia Tech
CAIA 2022
Water Infrastructure Security
Machine Learning Security
Attack Classification
Python
Scikit-learn
Pandas
NumPy
π Disclaimer
This project is intended for educational, research, and cybersecurity experimentation purposes.
The attack-detection component is designed to study the identification of malicious or abnormal behavior in controlled datasets and should not be interpreted as a complete security solution for real-world water infrastructure.
π Project Summary
Water Systems Attack Detection β CAIA 2022 is a machine-learning cybersecurity project focused on detecting abnormal and malicious behavior in water-system data.
The project combines:
Water-System Data
β
Data Preprocessing
β
Feature Engineering
β
Machine Learning
β
Attack Detection
β
Performance Evaluation
β
Security Analysis
The work demonstrates how machine learning can be used as a security layer for critical water infrastructure, while also highlighting the challenges of dataset dependency, class imbalance, false positives, and generalization.
π¨βπ» Modified by Momen
Original research/dataset context: CAIA 2022 / Virginia Tech Implementation: Python + Machine Learning Adaptation & Modifications: Momen