YAML Metadata Warning:empty or missing yaml metadata in repo card

Check out the documentation for more information.

Amazon Order Total Prediction & Segmentation

This project applies end-to-end data science techniques to an e-commerce (Amazon-like) dataset in order to:

1.Predict the final order total (TotalAmount)** using regression models.
2.Segment orders/customers into spending tiers** (low / mid / high) using classification.

The work includes data cleaning, exploratory data analysis (EDA), feature engineering, clustering, regression modeling, classification modeling, and model deployment.

  1. Dataset Description

Source: Synthetic Amazon-style orders dataset (Amazon.csv). Size: 100,000 rows. Input features (examples): Product: ProductID, ProductName, Category, Brand Pricing: UnitPrice, Discount, Tax, ShippingCost Quantity: Quantity Transaction info: OrderDate, PaymentMethod, OrderStatus, City, Country, SellerID

  • Target (Regression):
    • TotalAmount – final amount paid for the order.

The main research question: Can we accurately predict the final order total (TotalAmount) of an Amazon order using customer, product, pricing, and transaction-level features?**

  1. Exploratory Data Analysis (EDA)

Key steps:

Data cleaning:

  • Parsed OrderDate as datetime.
  • Created time-based features such as Week (ISO week number).
  • Checked for missing values, duplicated rows, and basic consistency.

Descriptive statistics:

  • UnitPrice:
    • Mean β‰ˆ 303, Std β‰ˆ 172, Min = 5, Max β‰ˆ 600.
    • Strong positive correlation with TotalAmount (~0.72).
  • Quantity:
    • Mostly between 1 and 5 units.
    • TotalAmount grows almost linearly with Quantity.

image

image

image

image

image

image

image

-Outlier analysis:

  • Some very high totals due to combination of high UnitPrice & Quantity.

  • Kept most outliers, since they represent genuine high-value orders, which are important for modeling.

  • Seasonality / Black Friday analysis:

    • Computed Week from OrderDate.
    • Found that Week 47 (Black Friday) and Week 48 (Cyber Monday) have slightly higher average TotalAmount than other weeks.
    • Week 47 shows the clearest spike relative to the global weekly average.
  1. Feature Engineering I added multiple engineered features to improve predictive power:

3.1 Time-based features

  • Week: ISO week number for each order.
  • IsHighSeason: 1 if Week in {47, 48} (Black Friday/Cyber Monday), else 0.

3.2 Price-related flags

  • IsExpensiveItem: 1 if UnitPrice above median unit price, else 0.

Note: To avoid data leakage, we did not include direct formulas that reconstruct TotalAmount (e.g., ValueBeforeTax = UnitPrice * Quantity * (1 βˆ’ Discount)) in the final models.

3.3 Clustering-based features (Unsupervised Learning)

I applied K-Means clustering on standardized numerical features:

  • Features used for clustering:

    • UnitPrice, Quantity, Discount, Tax, ShippingCost
  • Steps:

    1. Standardized the features with StandardScaler.
    2. Trained KMeans(n_clusters=5, random_state=42).
    3. Added:
      • ClusterID – the cluster assignment for each order.
      • DistanceToCentroid – Euclidean distance from each point to its cluster center.
  • PCA Visualization:

    • Reduced the clustering space to 2D using PCA (PCA1, PCA2) and plotted the clusters.
    • Clusters showed clear separation, corresponding roughly to:
      • High unit price & high quantity orders (high spenders).
      • Low unit price & low quantity orders (low spenders).
      • Orders with strong discounts.
      • Orders with higher taxes/shipping.

These cluster-derived features were then used as additional inputs to the regression and classification models.

image

image

  1. Regression Modeling (Predicting TotalAmount)

4.1 Baseline Linear Regression

  • Features used:
    • Numeric: UnitPrice, Quantity, Discount, Tax, ShippingCost, Week
    • Categorical: Category (one-hot encoded)
  • Train/Test Split: 80% / 20%, random_state=42
  • Performance:
    • RΒ² β‰ˆ 0.909
    • MAE β‰ˆ 166
    • RMSE β‰ˆ 217

Interpretation:
The baseline model already explains β‰ˆ91% of the variance in TotalAmount, mainly driven by Quantity, UnitPrice, and Tax. Category has only a minor effect.

4.2 Linear Regression with Feature Engineering & Clusters

Using:

  • All original numerical features
  • Week, IsHighSeason, IsExpensiveItem, ClusterID, DistanceToCentroid
  • One-hot encoded Category

Performance:

  • RΒ² β‰ˆ 0.920
  • MAE β‰ˆ 156
  • RMSE β‰ˆ 204

This shows a modest improvement over the baseline, primarily due to cluster-based segmentation and simple time-based flags.

4.3 Tree-Based Regression Models

I then trained two advanced models on the engineered dataset:

  1. Random Forest Regressor
  2. Gradient Boosting Regressor

Results (approximate):

  • Linear Regression (engineered):
    • RΒ² β‰ˆ 0.92, RMSE β‰ˆ 204
  • Random Forest Regressor:
    • RΒ² β‰ˆ 0.9999, RMSE β‰ˆ 6.6
  • Gradient Boosting Regressor:
    • RΒ² β‰ˆ 0.996, RMSE β‰ˆ 46.1

Because the underlying relationship between features and TotalAmount is nearly deterministic (a pricing formula applied consistently), tree-based models can almost perfectly reconstruct the function mapping inputs to target.

Winning Regression Model: Random Forest Regressor β€” near-perfect performance, robust to non-linearities and feature interactions.

5. Classification Task: Regression-to-Classification

I reframed the problem as a multi-class classification task, predicting spending tier class instead of continuous TotalAmount.

5.1 Target Transformation (Part 7.1)

I converted TotalAmount into three classes using quantile-based binning on the training set:

  • Class 0: bottom 33% (TotalAmount < 443.53)
  • Class 1: middle 33% (443.53 ≀ TotalAmount < 1088.52)
  • Class 2: top 33% (TotalAmount β‰₯ 1088.52)

This ensures balanced classes and has a meaningful business interpretation: low-, mid-, and high-spending orders.

5.2 Class Balance Check (Part 7.2)

Class distribution:

  • Train:
    • Class 0 β‰ˆ 33%
    • Class 1 β‰ˆ 33%
    • Class 2 β‰ˆ 34%
  • Test:
    • Very similar proportions (β‰ˆ33% each)

Since the classes are well-balanced, accuracy and macro F1 are both appropriate evaluation metrics.


6. Classification Models (Part 8)

I trained three different classifiers using the same engineered feature set:

  • Numeric: UnitPrice, Quantity, Discount, Tax, ShippingCost, Week, IsHighSeason, IsExpensiveItem, ClusterID, DistanceToCentroid
  • Categorical: Category (one-hot encoded)

All models used a ColumnTransformer + Pipeline for preprocessing and training.

6.1 Logistic Regression (Multinomial)

  • Reasonable performance as a baseline.
  • Main issues:
    • Frequently confuses the middle class (1) with lower (0) and higher (2) tiers.
  • Limitation:
    • Struggles to capture non-linear relationships in the data.

6.2 Random Forest Classifier

  • Best overall performance:
    • Very high accuracy and macro F1.
    • Confusion matrix shows almost perfect classification.
  • Interpretation:
    • Effectively captures complex patterns and interactions.
    • Leverages clustering-based features (ClusterID, DistanceToCentroid) and raw numeric inputs.

6.3 Gradient Boosting Classifier

  • Strong performance:
    • High precision and recall across classes.
    • Slightly more errors than Random Forest, mainly between neighboring classes (1 ↔ 2).
  • Still a very competitive model, but not the top performer.

Critical mistakes considered: In a business context, false negatives on the high-spending class (Class 2 β†’ Class 0/1) are more costly than false positives, because they represent missed opportunities for targeting valuable customers.

Winning Classification Model: Random Forest Classifier β€” highest accuracy, macro F1, and most balanced confusion matrix.

7. Exported Models

The following models were exported as .pkl files for deployment:

  • winning_model_random_forest.pkl
    β†’ Random Forest regression model (pipeline including preprocessing).
  • rf_classifier_model.pkl
    β†’ Random Forest classification model (pipeline including preprocessing).
Downloads last month

-

Downloads are not tracked for this model. How to track
Inference Providers NEW
This model isn't deployed by any Inference Provider. πŸ™‹ Ask for provider support