YAML Metadata Warning:empty or missing yaml metadata in repo card
Check out the documentation for more information.
Amazon Order Total Prediction & Segmentation
This project applies end-to-end data science techniques to an e-commerce (Amazon-like) dataset in order to:
1.Predict the final order total (TotalAmount)** using regression models.
2.Segment orders/customers into spending tiers** (low / mid / high) using classification.
The work includes data cleaning, exploratory data analysis (EDA), feature engineering, clustering, regression modeling, classification modeling, and model deployment.
- Dataset Description
Source: Synthetic Amazon-style orders dataset (Amazon.csv).
Size: 100,000 rows.
Input features (examples):
Product: ProductID, ProductName, Category, Brand
Pricing: UnitPrice, Discount, Tax, ShippingCost
Quantity: Quantity
Transaction info: OrderDate, PaymentMethod, OrderStatus, City, Country, SellerID
- Target (Regression):
TotalAmountβ final amount paid for the order.
The main research question:
Can we accurately predict the final order total (TotalAmount) of an Amazon order using customer, product, pricing, and transaction-level features?**
- Exploratory Data Analysis (EDA)
Key steps:
Data cleaning:
- Parsed
OrderDateas datetime. - Created time-based features such as
Week(ISO week number). - Checked for missing values, duplicated rows, and basic consistency.
Descriptive statistics:
UnitPrice:- Mean β 303, Std β 172, Min = 5, Max β 600.
- Strong positive correlation with
TotalAmount(~0.72).
Quantity:- Mostly between 1 and 5 units.
TotalAmountgrows almost linearly withQuantity.
-Outlier analysis:
Some very high totals due to combination of high
UnitPrice&Quantity.Kept most outliers, since they represent genuine high-value orders, which are important for modeling.
Seasonality / Black Friday analysis:
- Computed
WeekfromOrderDate. - Found that Week 47 (Black Friday) and Week 48 (Cyber Monday) have slightly higher average
TotalAmountthan other weeks. - Week 47 shows the clearest spike relative to the global weekly average.
- Computed
- Feature Engineering I added multiple engineered features to improve predictive power:
3.1 Time-based features
Week: ISO week number for each order.IsHighSeason: 1 ifWeekin {47, 48} (Black Friday/Cyber Monday), else 0.
3.2 Price-related flags
IsExpensiveItem: 1 ifUnitPriceabove median unit price, else 0.
Note: To avoid data leakage, we did not include direct formulas that reconstruct TotalAmount (e.g., ValueBeforeTax = UnitPrice * Quantity * (1 β Discount)) in the final models.
3.3 Clustering-based features (Unsupervised Learning)
I applied K-Means clustering on standardized numerical features:
Features used for clustering:
UnitPrice,Quantity,Discount,Tax,ShippingCost
Steps:
- Standardized the features with
StandardScaler. - Trained
KMeans(n_clusters=5, random_state=42). - Added:
ClusterIDβ the cluster assignment for each order.DistanceToCentroidβ Euclidean distance from each point to its cluster center.
- Standardized the features with
PCA Visualization:
- Reduced the clustering space to 2D using PCA (
PCA1,PCA2) and plotted the clusters. - Clusters showed clear separation, corresponding roughly to:
- High unit price & high quantity orders (high spenders).
- Low unit price & low quantity orders (low spenders).
- Orders with strong discounts.
- Orders with higher taxes/shipping.
- Reduced the clustering space to 2D using PCA (
These cluster-derived features were then used as additional inputs to the regression and classification models.
- Regression Modeling (Predicting
TotalAmount)
4.1 Baseline Linear Regression
- Features used:
- Numeric:
UnitPrice,Quantity,Discount,Tax,ShippingCost,Week - Categorical:
Category(one-hot encoded)
- Numeric:
- Train/Test Split: 80% / 20%,
random_state=42 - Performance:
- RΒ² β 0.909
- MAE β 166
- RMSE β 217
Interpretation:
The baseline model already explains β91% of the variance in TotalAmount, mainly driven by Quantity, UnitPrice, and Tax. Category has only a minor effect.
4.2 Linear Regression with Feature Engineering & Clusters
Using:
- All original numerical features
Week,IsHighSeason,IsExpensiveItem,ClusterID,DistanceToCentroid- One-hot encoded
Category
Performance:
- RΒ² β 0.920
- MAE β 156
- RMSE β 204
This shows a modest improvement over the baseline, primarily due to cluster-based segmentation and simple time-based flags.
4.3 Tree-Based Regression Models
I then trained two advanced models on the engineered dataset:
- Random Forest Regressor
- Gradient Boosting Regressor
Results (approximate):
- Linear Regression (engineered):
- RΒ² β 0.92, RMSE β 204
- Random Forest Regressor:
- RΒ² β 0.9999, RMSE β 6.6
- Gradient Boosting Regressor:
- RΒ² β 0.996, RMSE β 46.1
Because the underlying relationship between features and TotalAmount is nearly deterministic (a pricing formula applied consistently), tree-based models can almost perfectly reconstruct the function mapping inputs to target.
Winning Regression Model: Random Forest Regressor β near-perfect performance, robust to non-linearities and feature interactions.
5. Classification Task: Regression-to-Classification
I reframed the problem as a multi-class classification task, predicting spending tier class instead of continuous TotalAmount.
5.1 Target Transformation (Part 7.1)
I converted TotalAmount into three classes using quantile-based binning on the training set:
- Class 0: bottom 33% (
TotalAmount < 443.53) - Class 1: middle 33% (
443.53 β€ TotalAmount < 1088.52) - Class 2: top 33% (
TotalAmount β₯ 1088.52)
This ensures balanced classes and has a meaningful business interpretation: low-, mid-, and high-spending orders.
5.2 Class Balance Check (Part 7.2)
Class distribution:
- Train:
- Class 0 β 33%
- Class 1 β 33%
- Class 2 β 34%
- Test:
- Very similar proportions (β33% each)
Since the classes are well-balanced, accuracy and macro F1 are both appropriate evaluation metrics.
6. Classification Models (Part 8)
I trained three different classifiers using the same engineered feature set:
- Numeric:
UnitPrice,Quantity,Discount,Tax,ShippingCost,Week,IsHighSeason,IsExpensiveItem,ClusterID,DistanceToCentroid - Categorical:
Category(one-hot encoded)
All models used a ColumnTransformer + Pipeline for preprocessing and training.
6.1 Logistic Regression (Multinomial)
- Reasonable performance as a baseline.
- Main issues:
- Frequently confuses the middle class (1) with lower (0) and higher (2) tiers.
- Limitation:
- Struggles to capture non-linear relationships in the data.
6.2 Random Forest Classifier
- Best overall performance:
- Very high accuracy and macro F1.
- Confusion matrix shows almost perfect classification.
- Interpretation:
- Effectively captures complex patterns and interactions.
- Leverages clustering-based features (
ClusterID,DistanceToCentroid) and raw numeric inputs.
6.3 Gradient Boosting Classifier
- Strong performance:
- High precision and recall across classes.
- Slightly more errors than Random Forest, mainly between neighboring classes (1 β 2).
- Still a very competitive model, but not the top performer.
Critical mistakes considered: In a business context, false negatives on the high-spending class (Class 2 β Class 0/1) are more costly than false positives, because they represent missed opportunities for targeting valuable customers.
Winning Classification Model: Random Forest Classifier β highest accuracy, macro F1, and most balanced confusion matrix.
7. Exported Models
The following models were exported as .pkl files for deployment:
winning_model_random_forest.pkl
β Random Forest regression model (pipeline including preprocessing).rf_classifier_model.pkl
β Random Forest classification model (pipeline including preprocessing).








