YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
π RecoverAI β Intelligent Payment Recovery ML Pipeline
RecoverAI is an end-to-end Machine Learning pipeline that predicts whether a failed financial transaction can be recovered. It generates a synthetic hybrid dataset from real Kaggle transactions and trains a robust CatBoost model.
π Table of Contents
- Problem Statement
- Project Features
- ML Pipeline Workflow
- Data Engineering & Features
- Project Structure
- Quick Start Guide
- Model Performance
π― Problem Statement
Every day, millions of financial transactions fail due to network errors, insufficient funds, or timeouts. Knowing exactly which transactions have a high probability of recovery can save businesses millions in lost revenue.
RecoverAI tackles this by predicting β using a CatBoost model trained on 300,000 hybrid transactions β the likelihood of a transaction being successfully recovered within 72 hours.
β¨ Project Features
π οΈ Hybrid Data Engineering
- Automatically downloads real-world data (PaySim, Financial Transactions, Credit Card Fraud) from Kaggle.
- Engineers a comprehensive 360,000-row dataset containing 26 carefully crafted features.
- Imputes missing categories, temporal patterns (hour, day, weekend), and customer behaviors.
π§ Machine Learning Engine
- Uses CatBoostClassifier natively handling categorical features without the need for manual one-hot encoding.
- Incorporates early stopping with AUC-optimised evaluation.
- Outputs detailed evaluation metrics and feature importance rankings.
π Exploratory Data Analysis (EDA)
- Includes automated scripts (
eda_analysis.py,eda_fast.py) to visualize the distributions of amounts, categorical balances, and recovery rates.
π¬ ML Pipeline Workflow
graph TD
classDef file fill:#fff3e0,stroke:#e65100,stroke-width:2px;
classDef process fill:#e0f7fa,stroke:#006064,stroke-width:2px;
classDef output fill:#e8f5e9,stroke:#2e7d32,stroke-width:2px;
subgraph Phase 1: Data Gathering
DL["download_datasets.py"]:::process -->|Downloads| Kaggle["Raw Kaggle CSVs<br/>(PaySim, etc.)"]:::file
end
subgraph Phase 2: Feature Engineering
Kaggle --> Build["build_recoverai_dataset.py<br/>Engineers 26 Features"]:::process
Build --> TrainCSV["recoverai_training.csv<br/>(300k rows)"]:::file
Build --> ValCSV["recoverai_validation.csv<br/>(60k rows)"]:::file
end
subgraph Phase 3: Model Training
TrainCSV --> Train["train_catboost.py<br/>CatBoost Classifier"]:::process
ValCSV --> Train
Train --> Model["recoverai_catboost.cbm<br/>(Trained Model)"]:::output
Train --> Metrics["metrics.json & feature_importance.csv"]:::output
end
π Data Engineering & Features
The dataset builder script constructs 26 predictive features across 5 main categories:
| Category | Features |
|---|---|
| Transaction | amount, payment_method, error_reason, card_type, merchant_category, amount_bucket |
| Customer | customer_segment, customer_age, account_balance, customer_tenure_months, previous_failed_attempts |
| Behaviour | retry_count, risk_score, recovery_attempt_count, transaction_frequency_30d, time_since_last_failure_hr |
| Context | bank, region, device_type, channel, hour_of_day, day_of_week, is_weekend |
| Notifications | notification_sent, opt_out_notification, treatment_action |
Target Variable: recovered_within_72h (Binary Classification: 0 or 1).
π Project Structure
Revenue-AI-Tracker/
β
βββ π README.md β You are here
βββ π LICENSE β MIT
βββ π .gitignore
β
βββ π₯ download_datasets.py β Downloads Raw Kaggle CSVs
βββ π§ build_recoverai_dataset.py β Builds engineered train/val datasets
βββ π eda_fast.py β Fast terminal-based EDA
βββ π eda_analysis.py β Full EDA with graph visualisations
βββ π€ train_catboost.py β Model training & evaluation script
βββ π run_download.py β Wrapper to trigger dataset download
βββ π§ͺ ci_local_test.py β Local sanity testing script
β
βββ π requirements.txt (Optional)
βββ βοΈ docker-compose.yml β Environment setup
β‘ Quick Start Guide
Follow these steps to run the complete pipeline locally:
1. Setup Environment
git clone https://github.com/viRAJ357/Revenue-AI-Tracker.git
cd Revenue-AI-Tracker
# Install required python packages
pip install pandas numpy catboost scikit-learn
2. Download Kaggle Datasets
(Note: Requires a valid kaggle.json token configured in ~/.kaggle/)
python download_datasets.py
3. Generate the Dataset
This will merge the raw datasets and engineer the 360,000-row output files.
python build_recoverai_dataset.py
4. Train the Model
Train the CatBoost model. Once completed, it will save recoverai_catboost.cbm and performance metrics.
python train_catboost.py
5. Run Exploratory Data Analysis
python eda_fast.py
π Model Performance (Expected)
Actual performance may vary slightly based on random seed and dataset generation.
| Metric | Target Score |
|---|---|
| π― Accuracy | ~ 74.43% |
| π AUC-ROC | ~ 0.8207 |
| β Best Iteration | Approx. 160-200 / 500 |
| ποΈ Training Rows | 300,000 |
| π§ͺ Validation Rows | 60,000 |
| π’ Features | 26 |
Built with β€οΈ
RecoverAI β Turning failed transactions into recovered revenue.