CASE STUDY · MACHINE LEARNING · DISTRIBUTED COMPUTING
Customer churn prediction
across single-machine and parallel execution.
I built and compared churn-prediction pipelines using Decision Trees, Random Forests and Neural Networks, then re-ran the supervised workload through Ray to evaluate how parallel execution changes scalability, efficiency and reproducibility without confusing compute performance with model accuracy.
THE QUESTION
Can behavioural customer features predict churn, and what changes when execution is parallelised?
The project started from approximately one million online-retail transaction records and transformed them into customer-level behavioural features suitable for churn modelling.
The analytical goal was twofold: compare predictive behaviour across model classes and compare the same modelling logic across single-machine and Ray-based parallel execution environments.
FEATURE ENGINEERING
Raw transactions became customer-level behavioural signals.
- Frequency from distinct purchase occasions.
- Gross and net revenue.
- Customer lifetime and recency.
- Variety score from unique products purchased.
- Return rate.
- Purchase density.
- Spend per variety.
Recency was used to define churn through an inactivity threshold, then deliberately excluded from the predictive feature set to avoid target leakage.
MODEL COMPARISON
Interpretability, resilience and non-linearity were compared directly.
A Decision Tree provided an interpretable baseline. Random Forest reduced the instability and variance of a single tree through an ensemble of 500 trees. A feed-forward neural network with two hidden layers explored non-linear interactions across the behavioural features.
Across the project, Random Forest and the Neural Network outperformed the single Decision Tree in predictive terms, while the Decision Tree remained valuable for explanation and stakeholder communication.
PARALLEL EXECUTION
Ray was used as a task-parallel execution layer.
The distributed implementation used Ray remote functions so Random Forest and Neural Network training/prediction tasks could execute concurrently across available CPU resources rather than sequentially in one process.
This was task parallelism rather than full data-parallel distributed learning: the dataset itself was not partitioned across nodes. The important engineering result was that Ray improved the execution model and scalability of experimentation, while the relative predictive ranking of the models remained broadly stable.
EVALUATION
Accuracy was not treated as the only useful metric.
Performance was compared using accuracy, sensitivity, specificity, AUC and other classification measures. This made it possible to distinguish models that detected more churned customers from models that were more conservative about falsely flagging active customers.
The project also reinforced a key ML lesson: infrastructure can improve throughput and experimental scalability without automatically improving predictive quality.
BUSINESS INTERPRETATION
Churn emerged as a behavioural decline problem, not a random event.
Customer lifetime, frequency and net revenue repeatedly surfaced as important signals. The analysis supported the view that attrition is strongly connected to weakening engagement over time and that retention systems should prioritise behavioural change rather than isolated transactions.
WHAT THIS DEMONSTRATES
Machine learning plus systems thinking.
This work demonstrates end-to-end customer analytics, feature engineering, leakage prevention, supervised model comparison, evaluation trade-offs, neural-network experimentation and task-parallel ML execution with Ray.
It also shows the distinction between model performance and computational performance—an important consideration when moving ML workflows from local experimentation toward scalable production or HPC environments.