Logistics · ML engineering
E-Commerce Fulfillment & Supply Chain Optimization
Boosting Delivery Success to 98% and Saving 3,360+ Parcels Annually
Parcel split, carrier scoring, and in-transit risk detection for an operation shipping about 4,000 parcels a month. Delivery success rose from 91% to 98%.

Impact snapshot
| Metric | Result |
|---|---|
| Delivery Success Rate | Increased from 91% to 98% |
| Parcel Loss Prevention | ~280 parcels retained / month |
| Fulfillment Volume | 4,000 parcels processed / month |
Tech stack
Cloud & platform · AWS, Middle East (me-central-1)
AWS
Python
FastAPI
Docker
Airflow
scikit-learn
XGBoost
MLflow
Argo CD
GitLab
Compute: Amazon EKS. The production rules engine runs as a FastAPI service (26 replicas). The data-quality API runs on the same cluster in dev and prod. Images live in Amazon ECR; traffic goes through an Application Load Balancer with ACM certificates.
Delivery: Argo CD GitOps, GitLab CI (k8s_service for the data-quality API).
Data plane: Amazon Aurora MySQL 8 (read/write cluster), Amazon S3 for ML artifacts, AWS Secrets Manager, Amazon CloudWatch.
Orchestration: Apache Airflow on EKS, with ETL jobs launched as Kubernetes pods (parcel-level features, warehouse and carrier configuration, shipping-account health, point-of-entry features).
Decision services
Serving: Python 3.11, FastAPI, Uvicorn, Gunicorn, Docker.
Parcel split: Google OR-Tools (pywraplp) with the SCIP integer solver. Each line item is assigned to exactly one parcel; packed volume cannot exceed warehouse bin capacity; the objective is the minimum number of parcels. Cold and ambient goods are packed separately.
Delivery scoring: XGBoost binary classification (scikit-learn, pandas, NumPy), tracked and versioned in MLflow. The model scores warehouse, carrier, and shipping-method combinations on historical success, including recent success rates, volume, distance, and point of entry.
In-transit risk: A separate early-detection XGBoost pipeline for DHL, FedEx, and UPS, trained on carrier tracking events.
Flight context: FlightAware AeroAPI. A Python client pulls airport departures and historical flights by designator, registration, or FlightAware flight id, in seven-day windows, and normalizes scheduled, estimated, and actual out/off/on/in times into pandas. Credentials and database access come from AWS Secrets Manager and Aurora MySQL. This sits beside routing as lane context; it is not the component that produced the delivery-success lift.
Intake: Data Quality API, FastAPI on EKS. POST /data-push appends inbound JSON to data_quality_raw on the write database and leaves it unprocessed until a downstream job picks it up. GET /healthz checks that MySQL connection. Logs go to CloudWatch. It is an intake door into Aurora, not the live carrier score.
Operations UI: Clarity (Streamlit) for model performance, inference review, and shipping-account lifecycle.
The Challenge
An e-commerce fulfillment operation handling about 4,000 shipments a month was losing revenue to three coupled problems: orders split into more parcels than necessary, carrier and shipping-account choice made without a live success probability, and a 9% delivery failure rate. Extra splits raised shipping spend. Weak routing and late visibility into failures delayed fulfillment and eroded customer trust. International orders added a further choice — which point of entry and which flights those lanes depend on — and temperature-controlled goods had to stay within cold-chain capacity. Upstream systems also needed a controlled way to land raw operational payloads before they were cleaned and used for scoring.
The Solution
Minimum parcel split
For each order, the rules engine builds a bin-packing problem and solves it with OR-Tools SCIP. The solver returns the smallest set of parcels that still respects volume and separates cold from ambient items, which cuts avoidable multi-parcel fees.
Scored carrier and route selection
Airflow keeps a parcel-level feature store and shipping-account KPIs in Aurora MySQL. At decision time the rules engine loads the current MLflow model and scores candidate warehouse, carrier, and shipping-method options, including point-of-entry risk. The assignment is the option with the highest predicted delivery success, not a static carrier rule. Proposed models are compared with the live model before promotion.
Flight-aware lane context
The FlightAware client queries AeroAPI for departures from an origin airport and for historical flights on a given ident. Results are flattened to a time-ordered frame of scheduled versus actual movements, so lane reliability can be read from real flight history rather than from carrier name alone.
Early failure detection
A second model reads carrier tracking events in the first days after ship and flags parcels likely to fail while there is still time to intervene. Clarity exposes those scores, account health, and model drift to operations.
Raw data intake
Partners and internal jobs push JSON into the Data Quality API. Each payload is stored as a row in data_quality_raw and left unprocessed until a downstream job picks it up. The service is deployed with the rest of the platform (EKS, Argo CD, ECR) and reports health against the same Aurora write database the scoring path uses.
Production path
The scoring API, data-quality API, model registry, and ETL run on EKS, deployed by Argo CD from Git, with secrets and logs in AWS. FlightAware access is a Python integration against AeroAPI, using the same secrets and MySQL conventions as the ML pipelines.
The ROI & Results
Deliverability
Automated scoring, point-of-entry choice, and error prevention raised fulfillment success from 91% to 98%.
Revenue protection
About 280 packages a month (3,360 a year) that would previously have failed were retained.
Margin and capacity
Fewer unnecessary splits reduced per-order shipping fees, and the same decision path handles the monthly volume without manual carrier selection.