Does sentiment predict behaviour?
If a customer gives you a low score on a survey, are they about to leave? The CEO wanted to know, so I linked hundreds of thousands of survey responses to what customers actually did in the months either side. The first answer was wrong, and the reason it was wrong was the most useful part.
Situation
The business tracked its Net Promoter Score closely, from in-app surveys answered by hundreds of thousands of customers.
Task
The CEO asked whether NPS could do more than measure mood: could a low score flag a customer who was about to slow down or leave?
Action
Linked every survey response to the customer's behaviour from six months before to five months after, compared the score groups with each other rather than before with after, and looked at whole distributions, not averages.
Result
NPS alone is a weak churn signal: the score groups overlap heavily. I recommended keeping NPS as a sentiment measure and using customers' own behaviour for early warning.
The question
NPS (Net Promoter Score) asks customers how likely they are to recommend you, on a 0–10 scale. Scores of 0–6 are detractors, 7–8 are passives and 9–10 are promoters. The business tracked it closely, and the CEO asked whether it could do more than measure mood: could a detractor score flag a customer who is about to reduce activity or leave?
To test that, I took every completed in-app survey response over a half-year window and joined each one to the customer's monthly behaviour. That covered transaction count, transaction value, revenue, active days, days with a zero balance, and average balance. I cohorted customers by the month they answered, so "month 0" is always the survey month. Then I lined everyone up from month −6 to month +5. I split the results by business vs personal customers, by how they joined (through the agent/referral network or digitally), and by score group.
Charts use illustrative data rebuilt from the shape of the real results: values are indexed and lightly perturbed, and the real figures stay with the company.
The trap: everyone peaks in the survey month
The first cut looked dramatic. For personal customers, activity climbed steadily for six months, peaked in month 0, then fell for every group. Business customers showed the same shape, only flatter. Detractors fell, but so did promoters. Read naively, it says "answering a survey makes people transact less", which is obviously wrong.
The cause is how the sample is built. These are in-app surveys, so they only reach customers who opened the app that month. Every respondent was active in month 0 by construction, often unusually active. A customer at a temporary high tends to drift back towards their normal level afterwards. That is regression to the mean, and it produces a "decline" after month 0 whatever the customer thinks of you. The same selection also makes the months before look like growth.
The fix is to stop comparing before against after. I compared groups with each other, and the next step I recommended was a matched baseline:
- Groups with each other. All three groups share the same selection, so differences between them are informative even when the levels aren't.
- Each group against matched non-respondents (recommended next step). These would be customers who weren't surveyed but had near-identical behaviour in the six months up to and including the survey month. They share the same peak, so whatever happens to them afterwards is the baseline. This wasn't built. The chart's second view illustrates the idea with the average respondent as a stand-in, since every respondent shares the same month-0 selection.
What was left after the fix
Once the month-0 artefact is taken out, the story is much quieter. Promoters started higher and held more of their survey-month level. Detractors gave back more, most visibly among personal customers. The direction is what you'd expect. But the averages hide how little this separates individual customers, and the last couple of months rest on the earliest, smallest cohorts. The pattern held for business and personal customers, and for both acquisition channels, with differences in level but not in the conclusion.
A small average gap can still be useful if it's consistent. So I compared each customer's 60 days before the survey with the 60 days after, and looked at the whole distribution in each group, not just the averages. If NPS were a good early warning, detractors' outcomes would sit clearly below promoters'. They don't: the bands sit almost on top of each other, before and after.
Each bar spans the 10th to 90th percentile of customers in that group, with the darker band showing the middle half (25th to 75th) and the diamond the median; paler bars are the 60 days before the survey, stronger bars the 60 days after. Values are rebuilt from the real percentile tables, indexed so the median detractor before the survey is 100, on a log scale. Everyone's lower tail slips a little after the survey (the same selection effect as above). The slider is hypothetical: it pulls promoters' after-survey distribution up and detractors' down, to show what a genuinely predictive score would look like. At "as observed" the bands almost fully overlap.
What I recommended
NPS is a sentiment metric, not an early-warning signal
It is worth tracking for what it is: how customers feel, and what they complain about in the free-text answers. But a score from one moment, collected only from customers active that month, separates future behaviour too weakly to target retention work. For "who is about to slip away", the better input is behaviour itself: each customer's recent activity against their own history. That is the approach I took in an early-warning score for top business customers, piloted in one region.
Two practical changes followed from the analysis. Any report that puts NPS next to behaviour should use a matched baseline, never a before-and-after line. And if survey scores feed into a risk model at all, they should be one feature among behavioural ones, tested for whether they add anything.
Caveats and what I'd do differently
- Matching is only as good as what you match on. The matched baseline I recommended would match on past activity. Customers who choose to answer a survey may differ in ways activity doesn't capture, such as engagement with notifications, so some selection would likely remain.
- Later cohorts have shorter follow-up. Customers surveyed late in the window only have a month or two of "after". I'd report results by cohort or restrict to complete follow-up windows rather than pooling.
- Monthly grain blurs timing. A customer surveyed on the 1st and one surveyed on the 30th are both "month 0". Daily data aligned to the exact response date would sharpen the picture.
- Score changes may carry more signal than levels. A customer who drops from promoter to detractor between surveys is a different case from a habitual detractor. I'd test that with repeat respondents.
- Use the text. The free-text reasons are likely more actionable than the number, especially when they mention a specific failure such as a declined transfer.