Problem
This portfolio exercise explores how transaction history can be grouped into descriptive customer segments. It is not an analysis commissioned by or performed for an actual retailer.
Data
The repository identifies its Australian retail dataset as synthetic and for demonstration only. The notebook groups transactions by customer and derives recency from the latest order date, frequency from unique order counts, and monetary value from summed net revenue.
Approach
The notebook assigns 1–5 scores to recency, frequency, and monetary value using quantile-based bins, then applies explicit rules to assign labels such as Champions, At Risk, and Potential Loyalists. It also creates segment summaries and visualizations.
Technology
Python, pandas, NumPy, Matplotlib, and Seaborn.
Solution
A Jupyter notebook documents the customer-level aggregation, scoring, segment assignment, and summary charts. The repository includes visuals for segment distribution, RFM distributions, and revenue by segment and state.
Sample results and limitations
The notebook produces rule-based customer groups from the synthetic transactions. These labels and totals describe only the sample; they do not establish actual customer churn, lost revenue, recoverable revenue, or the effect of a marketing campaign.
Methodological learning
RFM is a descriptive segmentation approach. The chosen scoring thresholds and segment rules shape the resulting groups, so labels such as “At Risk” should not be treated as predictions of real customer behavior without further validation.