Prosper Loan Analysis

Colab Notebook

https://colab.research.google.com/drive/10Rt5BJyX-tcX8mHaiFN_KVzcPWG23YYk#scrollTo=w84cR3AZIU0e

Part 1 - Dataset Overview

Dataset Description

The dataset used in this project is the Prosper Loan Dataset, obtained from Kaggle. Prosper is a U.S. peer to peer lending platform where individuals can apply for and invest in personal loans. The dataset contains detailed borrower-level financial, demographic, credit, and loan-performance information collected throughout the loan application and servicing process.

Each record represents a single loan, described through a structured set of borrower characteristics, credit history indicators, and loan attributes.

Borrower & Demographic Attributes- Employment Status, Occupation, Stated Monthly Income, Employment Length.

Credit-Related Features- Credit Score Range, Number of Credit Lines, Delinquency History, Public Records, Revolving Credit Balance, Bankcard Utilization, and other creditworthiness indicators.

Loan Characteristics- Loan Amount, Loan Term, BorrowerAPR (effective interest rate), Monthly Payment, Percent Funded, Number of Investors, Prosper Rating, Origination Date.

Economic & Performance Indicators- Debt to Income Ratio, Collection Fees, Principal Payments, Loss Metrics recorded during the loan lifecycle.

Objective of the Analysis

The goal of this analysis is to identify which borrower, credit and loan attributes influence the effective interest rate (BorrowerAPR) in the Prosper lending platform. By examining variations in income levels, credit scores, debt to income ratios, employment status and credit history indicators, we aim to uncover the key factors that shape risk based loan pricing.

The analysis focuses on how core financial and demographic attributes such as stated monthly income, credit score ranges, debt to income ratio, revolving credit balance, delinquency history and credit history length reflect borrower stability and creditworthiness. Understanding these relationships helps explain the underlying drivers of interest rate determination.

Target Variable

The target variable in this analysis is BorrowerAPR, the effective interest rate assigned to each loan and a direct reflection of borrower risk. The goal is to understand which financial and credit factors drive APR higher or lower. Later, this continuous variable is also grouped into ranges for classification, allowing us to explore how borrower characteristics map to different APR levels.

Part 2 - Exploratory Data Analysis

Data Cleaning

The Prosper dataset underwent a comprehensive cleaning process to ensure high data quality and prevent leakage in later modeling. Columns with more than 70 percent missing values were removed because they provided limited usable information and could not be reliably imputed. This included outdated or sparsely populated credit grade fields. Additional variables such as ClosedDate and Prosper’s internal model outputs were dropped because they reveal information from after the APR decision and would introduce leakage into the analysis.

Columns with moderate missingness, including Prosper rating fields and estimated performance metrics, were retained for exploration since their relevance had not yet been fully evaluated. Missing categorical values were filled with an explicit “Unknown” label, while numerical fields were imputed using the median to preserve distributional integrity in the presence of skewed financial data. Features with very small fractions of missing values followed the same strategy, ensuring completeness without distorting statistical patterns.

We confirmed that no duplicate records existed, and key categorical fields were standardized by trimming spaces and normalizing text formats. Sanity checks were performed to detect unrealistic values. Structurally impossible entries such as percent funded values above one and excessively large employment duration values were removed, while extreme but plausible financial values such as high revolving credit balances were preserved for later outlier analysis.

Outlier Detection & Handling

The initial sanity checks guided which variables required focused outlier inspection, leading us to examine stated monthly income, debt to income ratio and revolving credit balance. To improve readability, some outlier plots use a capped or broken x axis. This helps display the main distribution clearly while still showing extreme values, without altering the underlying data.

For stated monthly income, most extreme values were legitimate high earners and were retained. A sanity threshold of 250,000 was applied to flag implausible values, and only 6 records were removed. For debt to income ratio, statistical outliers began around 0.8, but values up to 3 are still economically reasonable, so they were kept. Only ratios above 3 were considered unrealistic and removed. For revolving credit balance, most borrowers fell within a normal range, but a threshold of 300,000 identified values unlikely to reflect true consumer credit behavior. Less than 0.2 % of records exceeded this limit and were removed.

These steps allowed us to preserve meaningful financial variation while eliminating values that were clearly implausible or indicative of data errors.

Across the three features examined for outliers (StatedMonthlyIncome, DebtToIncomeRatio and RevolvingCreditBalance), the IQR method flagged many observations as statistical outliers, though most represented legitimate financial variation. Instead of removing all flagged values, we applied domain based sanity thresholds to distinguish plausible extremes from economically impossible entries. Only 610 records (less than 0.61 % of the dataset) exceeded these limits and were removed. This selective approach preserves meaningful high value cases while improving overall data quality.

Descriptive statistics summary

•	Borrowers earn 4600-6800 per month and have 5-6 years of employment, indicating generally stable income.
•	Median revolving payments (271) and balances (8500) show typical consumer credit use.
•	Bankcard utilization averages 60%, reflecting moderate reliance on available credit.
•	Borrowers hold 9–10 active credit lines, about 25 total lines, and 6 revolving accounts, showing consistent engagement with credit products.
•	BorrowerAPR is usually 21%, interest rates 18–19%, and most loans are 36 months, with purposes concentrated in debt consolidation, home improvement and business.
•	Most borrowers have 0–1 recent inquiries, 4–7 total inquiries, and rare delinquencies, indicating clean credit histories.
•	Borrowers with stable employment receive lower APRs, while those with less reliable income face higher rates due to increased perceived risk.

Vizualizations

To keep the README concise, only a selected subset of visualizations is included here. The full analysis contains additional plots and deeper exploratory visuals, all of which are available in the accompanying Colab notebook.

Distribution of BorrowerAPR

The distribution of BorrowerAPR shows that most loans fall within the 15%–25% range, with a smaller right-skewed tail representing higher risk borrowers who receive significantly higher rates. This pattern is typical in consumer lending and reflects how lenders price loans based on perceived borrower risk.

Top Loan Purpose Categories

This chart displays the most common loan purposes among borrowers, using Prosper’s official category definitions. Debt Consolidation dominates the platform, accounting for over half of all loans, followed by categories such as Other, Not Available, Home Improvement and Business. The distribution highlights the primary financial motivations driving borrowers to seek funding.

Relationships Between Interest Rates, Loan Size, Payments, and Investor Activity

The heatmap reveals strong correlations among BorrowerAPR, BorrowerRate and LenderYield, which represent different forms of the interest rate. LoanOriginalAmount is closely tied to MonthlyLoanPayment, as larger loans require larger payments. Investor count shows moderate positive correlation with loan size and a negative relationship with APR, indicating that larger, lower risk loans attract more investors. Overall, the heatmap illustrates the structural relationships that shape loan pricing and borrower risk.

Research

Only a subset of research visualizations is shown here for readability. Full analysis with all six research questions and plots is available in the Colab notebook.

How Does Employment Status Influence Borrower APR?

Employment stability strongly influences APR: stable workers receive the lowest rates, while self employed and unemployed borrowers pay substantially more due to higher perceived risk. This raises the question of whether income level can further explain these differences, specifically, how income interacts with employment status to shape APR.

How Does Income Modify the Effect of Employment Status on APR?

Higher debt to income ratios are associated with higher APRs, showing that borrowers with heavier debt loads are priced as riskier. Income further intensifies this pattern: high income borrowers receive lower APRs, while low income borrowers pay more even at similar DTI levels. Together, DTI and income act as independent signals in lenders' risk pricing. This naturally leads to examining another core dimension of borrower risk: credit score.

How does a borrower’s credit score relate to their APR?

Average APR decreases consistently as credit scores improve, showing that borrowers with weaker credit profiles face substantially higher APRs, while those with stronger scores receive much lower APRs. This underscores credit score as a central signal of repayment reliability in the underwriting process. To extend this analysis, we next examine whether past delinquency behavior offers additional explanatory power beyond credit score alone.

How do borrowers' past credit delinquencies influence the APR assigned to them?

Average APR increases steadily as the number of past delinquencies rises, indicating that lenders strongly penalize repeated late payments. Borrowers with clean repayment histories receive the lowest APRs, while those with multiple delinquencies face substantially higher APRs. This highlights past repayment behavior as a key driver in APR assignment.

Conclusions From Our Research

Across all analyses, lenders appear to price APR using multiple signals of borrower risk. Lower income, smaller loans, higher DTI, and prior delinquencies all correspond to higher APRs. Credit score, approximated from each borrower's reported range, shows a strong inverse relationship with APR. Together, these factors form a consistent risk framework in which APR rises as financial stability decreases.

Part 3: Define and Train a baseline model

Regression Goal:

The goal of this regression task is to predict a borrower's APR based on financial, behavioral, and credit-related features available at the time the loan was issued.

Configuring the Features for Baseline Modeling

For the baseline regression model, we selected only numerical features available at the time of loan origination to create a simple, unbiased starting point. Limiting the model to numeric variables avoids encoding steps and reflects the raw predictive signal in the data. We excluded identifiers, dates, categorical fields, and any post loan performance information to prevent leakage and ensure that all inputs represent information truly known when the APR was assigned.

Train Test Split and Baseline Model Training

We defined BorrowerAPR as the target variable (y) and the selected numerical predictors as the feature matrix (X). The data was then split into training and testing sets, reserving 20% for unbiased evaluation while ensuring reproducibility with a fixed random seed. A simple Linear Regression model was trained on the training set to form a baseline, allowing us to assess the inherent predictive signal in the numerical features before applying more advanced modeling techniques.

Baseline Model Evaluation

We evaluated the baseline Linear Regression model using MAE, MSE, RMSE, and R². RMSE was used as the primary metric because it penalizes large errors and is expressed in the same units as APR, making it intuitive to interpret. To contextualize model performance, we compared the RMSE (0.0615) with the actual distribution of BorrowerAPR values. Given that most APRs fall between 0.15 and 0.28, this error level is reasonable for an unengineered baseline and highlights clear room for improvement through feature engineering and more advanced models.

MAE: 0.0487 | MSE: 0.0038 | RMSE: 0.0615 | R²: 0.4115

Actual vs. Predicted APR

The scatter plot shows wide deviations between actual and predicted APRs, with over-prediction at low APRs and under-prediction at high APRs. This spread around the perfect-fit line confirms that the baseline model captures only part of the signal and requires further improvement.

Residuals vs. Predicted Values

The residual plot shows a clear downward trend, meaning the model overestimates low APRs and underestimates high APRs instead of producing random residuals around zero. This systematic pattern, along with the widening spread, indicates missed non linear relationships and heteroscedasticity that limit the baseline model’s accuracy.

Feature Importance

The baseline linear model highlights several intuitive relationships. Features associated with higher risk such as Debt to Income Ratio and Bankcard Utilization receive positive coefficients, increasing predicted APR. In contrast, strong credit signals such as a higher percentage of never delinquent trades receive negative coefficients, reducing APR.

Although the model is simple and cannot capture non linear patterns, its coefficient directions align well with financial logic. This provides a clear baseline understanding of which borrower characteristics most strongly influence APR before moving on to more advanced models.

Part 4: Feature Engineering

Removing Features

We removed additional columns that cause data leakage because they are generated after APR is determined or depend on Prosper’s internal scoring model. We also removed identifier columns that provide no predictive value. These features were not used in the baseline model, and excluding them ensures a cleaner and leakage free dataset for training the advanced models.

New Features

To strengthen the model, we engineered several features that capture combined financial behaviors and relationships not visible from individual variables alone.

•	CreditHistoryYears - measures the length of each borrower's credit history at the time of loan origination.
•	CreditScore - midpoint of the reported lower and upper credit score bounds, creating a single interpretable score.
•	IncomeStabilityRatio - compares monthly income to debt burden, indicating repayment capacity.
•	RevolvingBalanceBurden - reflects how heavy revolving balances are relative to income, signaling financial strain.
•	DelinquencyPressureIndex - combines long term and current delinquencies into a stronger risk indicator.
•	LoanToIncomeRatio - quantifies how large the requested loan is relative to income, capturing repayment pressure.

This compact feature set provides the model with richer risk-related signals while remaining fully interpretable.

Encoding

To prepare categorical data for modeling, we converted all non-numeric fields into numerical representations. Boolean variables were mapped to 0/1, meaningful date components were extracted from loan timing fields, and one hot encoding was applied to categories such as employment status and loan quarter.

For high cardinality features, we used frequency encoding to avoid creating excessive columns. These transformations preserve the information in each variable while producing a clean, model ready numerical dataset.

Scaling Numerical Features (Standardization)

The numerical features in the dataset span very different scales (for example, income values in the thousands vs. ratios between 0 and 1). To ensure that all variables contribute equally during training, we apply StandardScaler, which standardizes each feature to have mean 0 and standard deviation 1. This prevents large scale features from dominating the model and helps regression algorithms converge more effectively.

Applying Borrower Risk Clustering (Unsupervised Learning)

In this section, we use K-Means clustering to group borrowers into behavioral risk segments.
The clustering is based on standardized financial features that capture credit usage, delinquency history, and overall credit activity, so that each cluster reflects a distinct borrower risk profile.

Elbow Method: Determining the Optimal Number of Clusters

We evaluated how WCSS changes across different values of k. The curve shows a sharp improvement up to k = 3, after which gains diminish sharply. This indicates that three clusters capture most of the structure in the borrowers’ financial behavior without adding unnecessary complexity.

Silhouette Analysis: A Second Method for Choosing the Optimal Number of Clusters

We computed silhouette scores for the same range of k. The highest score also appears at k = 3, indicating that this value produces the most coherent and well separated groups. Conclusion: Three clusters form the most stable and meaningful segmentation.

Based on the Elbow and Silhouette results, we fit K-Means with k = 3.

PCA Visualization of Borrower Risk Clusters

This PCA scatter plot visualizes the three K-Means borrower clusters after reducing the variables to two components for clarity. The axes themselves have no financial meaning, but the separation between clusters is clear. Each group reflects a distinct financial pattern: one shows clean repayment behavior and low credit utilization, another exhibits higher delinquency and credit pressure, and the third displays moderate, mixed behavior. These consistent differences indicate that the clusters capture real borrower profiles rather than random groupings. The resulting RiskCluster feature therefore provides a compact and meaningful representation of borrower risk for downstream modeling.
Feature Cluster 0 Cluster 1 Cluster 2
CreditScore 0.078 -0.805 0.227
DebtToIncomeRatio -0.072 -0.405 0.355
DelinquenciesLast7Years -0.272 1.754 -0.274
RevolvingCreditBalance -0.154 -0.431 0.545
TradesNeverDelinquent (%) 0.17 -1.61 0.41

To understand what differentiates the clusters, we examine the mean standardized feature values for each group. Cluster 0 shows average and stable financial behavior, Cluster 1 represents borrowers with the weakest credit indicators and highest delinquency levels, and Cluster 2 reflects higher income and activity but low delinquencies. These numerical patterns confirm the PCA separation and validate that the clusters represent meaningful borrower profiles.

In addition, Each borrower now gets three numbers measuring how close they are to each cluster's centroid. Smaller distance means more similar behavior. These continuous signals give the model richer information than a single cluster label.

Re-Scaling the Dataset After Adding New Features

We applied scaling again after creating the new cluster based features because their numeric ranges differ from the existing variables. Re-scaling ensures that all features operate on a comparable scale and prevents the model from unintentionally overweighting the newly added features. We exclude RiskCluster from scaling because it is a categorical group label (0/1/2), treating it as a continuous value would distort its meaning and mislead the model. Rescaling at the final stage of feature engineering is standard and ensures a clean, model ready dataset.

Part 5+6: Regression Models

Linear Regression again

Linear Regression shows the same performance as the baseline because scaling and feature engineering do not add new linear signal for it to learn. Since the model can only capture linear relationships, its predictions remain unchanged, indicating that more advanced models are needed to benefit from the engineered features.

Ridge Regression Model with GridSearchCV

Ridge Regression significantly improves over Linear Regression: all errors decrease (MAE, MSE, RMSE) and R² rises from 0.42 to 0.53. By shrinking coefficients and handling multicollinearity, Ridge makes better use of the engineered and cluster-based features, resulting in a more stable and accurate model.

MAE: 0.0436 | MSE: 0.0030 | RMSE: 0.0549 | R²: 0.5311

Ridge Regression - Actual vs. Predicted APR Visualization

The Ridge model's predictions align much more closely with the diagonal line, showing reduced scatter and tighter clustering around the true APR values. This indicates that Ridge produces more stable and accurate predictions than Linear Regression, especially across the mid-range of APRs.

Ridge Regression - Residuals vs. Predicted APR Analysis

The Ridge residuals are much more tightly clustered around zero, with far less visible trend compared to Linear Regression. This indicates reduced bias and more consistent errors across the entire APR range, showing that Ridge provides noticeably more stable and reliable predictions.

Feature Importance

The plot shows the top 8 most influential features in the Ridge model. Positive coefficients (green) increase predicted APR, while negative coefficients (red) decrease it. Distance to cluster features and key financial indicators like Debt to Income Ratio and loan year effects are among the strongest contributors, showing that both behavioral risk and financial structure drive the model's predictions.

Conclusion of Ridge

Ridge Regression delivers a clear improvement over the baseline model by reducing errors and providing more stable, reliable predictions. Its regularization helps handle multicollinearity, allowing the model to better leverage the engineered features.

XGBoost Regression Model

The XGBoost model delivers a substantial performance boost, achieving much lower errors and a high R² of 0.75. This indicates that XGBoost captures non-linear relationships and complex interactions that linear models cannot. Overall, it provides the strongest and most accurate APR predictions among all tested models.

MAE: 0.0300 | MSE: 0.0016 | RMSE: 0.0401 | R²: 0.7502

XGBoost Regression - Actual vs. Predicted APR Visualization

The XGBoost predictions align much more closely with the diagonal line, showing far greater accuracy than Linear Regression and Ridge. Systematic errors are largely reduced, and the scatter is tighter and more consistent, reflecting XGBoost's ability to capture nonlinear relationships that linear models miss.

XGBoost Regression - Residuals vs. Predicted APR Analysis

The XGBoost residual plot shows residuals tightly centered around zero with no strong pattern, indicating low bias and more stable errors across all APR levels. Compared to Ridge, the systematic under and over prediction patterns largely disappear, reflecting XGBoost's ability to capture nonlinear relationships and produce more reliable predictions.

Feature Importance

XGBoost uncovers a distinct pattern of feature importance, highlighting relationships that linear models cannot capture. The model ranks CreditScore as the most influential factor, followed closely by several LoanYear features, suggesting that historical market conditions and lending environments significantly shape APR outcomes. Features like AvailableBankcardCredit reflect borrowers' available credit and financial stability, further contributing meaningful predictive value. Overall, XGBoost identifies nonlinear interactions and temporal effects that traditional linear models tend to overlook.

Summary and Conclusions

Model MAE MSE RMSE R²
Linear Regression 0.0485 0.0037 0.0615 0.418
Ridge Regression 0.0436 0.0030 0.0549 0.531
XGBoost Regression 0.0300 0.0016 0.0401 0.750

XGBoost delivers by far the best performance among all models tested. Unlike Linear Regression and Ridge, which captured only simple linear trends, XGBoost learned nonlinear relationships and feature interactions, achieving much higher accuracy (R² = 0.75) and significantly lower prediction errors.

Its strong results were powered by the earlier feature engineering: credit behavior features, temporal loan indicators, and cluster based risk variables became highly influential, enabling the model to detect complex borrower risk patterns.

Overall, the combination of rich engineered features and a nonlinear model makes XGBoost the most effective and reliable predictor of BorrowerAPR.

Part 7+8: Classification Models

We converted APR into three classes using quantile binning to create a balanced distribution across low, medium, and high APR levels. This preserves the natural ordering of APR while avoiding class imbalance, ensuring that classification models learn meaningful risk tiers.

Logistic Regression

Random Forest Classifier

XGBoost Classifier

Summary and Conclusions

We trained three classification models - Logistic Regression, Random Forest, and XGBoost and evaluated their performance using confusion matrices to understand how well they distinguish between the three APR classes (low, medium, high).

Across all three models, XGBoost delivers the most accurate and consistent predictions. Logistic Regression struggles with non linear patterns and often confuses neighboring APR classes. Random Forest improves class separation but still shows noticeable misclassification in boundary cases. XGBoost achieves the clearest diagonal pattern in the confusion matrix, indicating far fewer false negatives and false positives, and demonstrates the strongest ability to capture complex relationships learned during feature engineering.

Overall, XGBoost is the best performing classifier, offering the highest reliability for distinguishing APR risk tiers.


Extra Work (short version):

Higher APR segments appear strongly associated with instability indicators such as high DTI and delinquency pressure. Such patterns highlight the ethical consideration that financially stressed borrowers often pay disproportionately higher rates. Models predicting APR must therefore be used responsibly to avoid reinforcing disadvantage.

Key Takeaways

Through iterative training, I observed how significantly model performance improves when the features are meaningful. The linear models showed only limited gains, but after adding richer engineered features—temporal signals, credit behavior variables, and clustering, the more advanced models, especially XGBoost, improved dramatically. This reinforced the idea that strong feature engineering often matters more than the choice of algorithm.

Challenges and Key Lessons Learned

I initially began the project with a different dataset and invested quite a lot of work exploring and preparing it, but eventually realized it wouldn’t support the type of analysis I wanted, so I decided to switch. Later in the process, I also learned an important pipeline lesson: after accidentally overwriting the train/test split following feature engineering, all model performance dropped, and it took hours to trace the issue. This reinforced how essential clean, consistent workflow steps are for reliable modeling.

Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. 🙋 Ask for provider support