A health score is not a measurement. It is a prediction, and an unvalidated prediction is a decoration.
The question that separates a useful score from a coloured dot: of the customers who churned last year, what did the score say ninety days before? If nobody has run that check, the score is not doing the job it exists for.
Behaviour beats sentiment
Inputs that predict, roughly in order:
- Usage depth and trend. Not logins. Which features, how often, by how
many people, and which direction it is moving.
- Breadth of adoption. One user is a single point of failure; twenty is a
habit.
- Champion status. Did the person who bought it still work there last
month? Departures are among the strongest available churn signals.
- Support pattern. Not ticket volume, which is ambiguous. Unresolved
tickets and escalations.
- Commercial signals. Late payment, downgrade requests, procurement asking
about terms.
Sentiment inputs such as CSAT and NPS belong in the model at low weight. They come from people who responded, which is a biased sample of the ones you most need to predict.
Validate it, then keep validating
Take the customers who churned in the last twelve months. Pull their health score ninety days before churn. If most were green, the model is wrong.
Two numbers to track, borrowed from any classifier:
- How many churned customers the score flagged in advance. Missing them is
the expensive failure.
- How many flagged customers actually churned. Flagging everyone catches
every churn and is useless.
Re-run quarterly. Product changes and packaging changes both invalidate weights that were fine last year.
The failure mode
A weighted average of six inputs, weights chosen in a workshop, rendered as red, amber, green, never checked. It produces a dashboard everyone trusts and nothing predictive, and it fails most badly at exactly the moment it matters, because the inputs that would have caught a quiet non-renewal were never in it.
The specific gap: most scores measure engagement with the product and ignore whether the person who championed it still works there.
A worked example
A model weighting usage trend 40%, adoption breadth 25%, champion present 20%, unresolved support 10%, sentiment 5%.
Back-test: 34 customers churned last year. 90 days out, 21 were flagged amber or red and 13 were green. That is 62% caught. Of 88 accounts flagged in the same period, 21 churned, so 24% of flags were real.
Both numbers are usable and neither is good. The 13 missed accounts are the place to look: if most lost their champion, the champion weight is too low.
The traps
- Never back-testing.
- Sentiment weighted heavily.
- Login counts as the usage input.
- A single score across segments that behave differently.
- Fixed weights after a packaging change.
Where this sits
Health scoring is Measurement in the Tenbound Pipeline Architecture Standard, and the strongest input belongs to Signal: a champion leaving is the same class of event that makes an account worth calling in outbound. The difference is which direction you act.