Credit Risk Prediction using Ensemble Learning

This project aims to predict the probability of credit default using a combination of Logistic Regression, Random Forest, and XGBoost classifiers combined in a soft voting ensemble. The dataset contains financial, demographic, and credit history attributes of loan applicants. The objective is to identify high-risk applicants and support better lending decisions.

Dataset

Source: Give Me Some Credit.csv Target variable: SeriousDlqin2yrs

  • 0 β†’ No serious delinquency in the past 2 years
  • 1 β†’ Serious delinquency occurred

Key Features:

  • RevolvingUtilizationOfUnsecuredLines – Ratio of credit card balance to credit limit
  • NumberOfTime30-59DaysPastDueNotWorse – Count of 30-59 days late payments
  • age – Age of the applicant
  • NumberOfTimes90DaysLate – Count of 90+ days late payments
  • DebtRatio – Monthly debt payments to income ratio
  • MonthlyIncome – Monthly income of the applicant
  • NumberOfOpenCreditLinesAndLoans – Number of open credit lines and loans
  • ...and other relevant credit-related attributes.

Data Preprocessing

  • Handling Missing Values: Median imputation for numerical columns.

  • Outlier Treatment: Clipping based on IQR method.

  • Feature Engineering:

    • DebtToIncomeRatio = RevolvingUtilizationOfUnsecuredLines / MonthlyIncome
    • Age binning into categories (<30, 30-40, 40-50, 50-60, 60+)
  • Scaling: StandardScaler applied to numerical features.

  • Class Imbalance Handling: Stratified train-test split to maintain target distribution.

Model Training

Three base models were trained and tuned using GridSearchCV with ROC-AUC as the scoring metric:

  1. Logistic Regression (with L1 regularization)
  2. Random Forest Classifier
  3. XGBoost Classifier

These were combined into a VotingClassifier with voting="soft" to leverage predicted probabilities from all models.

Model Performance

Metric Class 0 Class 1
Precision 0.94 0.60
Recall 0.99 0.18
F1-Score 0.97 0.28
Accuracy 0.94
ROC-AUC Score 0.864

Macro Avg F1: 0.62 Weighted Avg F1: 0.92

Feature Importance (SHAP Analysis)

The top features influencing the model’s predictions are:

  1. RevolvingUtilizationOfUnsecuredLines
  2. NumberOfTime30-59DaysPastDueNotWorse
  3. Age
  4. NumberOfTimes90DaysLate
  5. NumberOfOpenCreditLinesAndLoans
  6. DebtToIncomeRatio

πŸ“Œ Conclusion

  • The ensemble approach achieved a ROC-AUC of ~0.864, showing strong discriminatory power.
  • SHAP analysis highlighted that credit utilization, late payment history, and age are key determinants of credit risk.
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support