Customer satisfaction metrics: how one question became CSAT, NPS and CES

The main customer satisfaction metrics are CSAT, NPS, and CES. See what CSAT score is, how to calculate it, what a good score looks like, and how to act on it.

Gourab Majumder
Last updated August 19, 2026 13 min read
A young child laughing, photographed close up in black and white, with a pale flower held up beside their face.

Customer satisfaction metrics all descend from one question: how satisfied were you. CSAT still asks it. NPS and CES were built on arguments that it was not enough, and most published advice now treats the three as rivals to be ranked by predictive power. This article is about satisfaction itself: where measuring it began, why satisfied turns out to be a weak state, where the satisfaction question still belongs, and how the metrics carved out of it sit at different layers of the customer relationship.

TL;DR

  • Satisfaction was the first thing anyone measured in service feedback. NPS and CES were both created on arguments that satisfaction was failing at part of its job; when NPS's founding claim was later tested, the two measures came out level.
  • Satisfied is a weak word on its own terms: it tells you nothing went wrong, not that the customer was delighted or got what they came for.
  • The three metrics sit at three layers. NPS is the relationship layer, CSAT is the journey layer, CES is the process layer. Choose by layer, not by predictive claims.
  • CSAT is a genuine percentage. NPS is a points score from -100 to +100 and is never a percentage. Putting the two side by side without saying so compares different kinds of number.
  • There is no universally good CSAT score: the level of any score depends on category, collection mode and who chose to answer. Numr publishes no benchmark, as a position.

Satisfaction came first, and the others were carved out of it

Satisfaction was the first thing anyone measured in service feedback. It is why these are still called satisfaction surveys: "how satisfied were you" is literally what was asked. The American Customer Satisfaction Index measured it at national scale.

Then the carving started. NPS was created on the argument that satisfaction did not correlate with anything financially driven. The recommendation question, the question NPS is built on, was offered as reaching for a stronger state than satisfaction. That founding argument, it is worth saying in passing, was a claim rather than a finding: when a 2007 Journal of Marketing study tested NPS and a national satisfaction index against firm revenue growth, the two came out level. CES followed on a different complaint: satisfaction did not tell you whether the customer could get their job done. A customer can report being satisfied with an interaction that took far longer to resolve than it should have, and the satisfaction score will not surface the struggle.

Each new metric took a piece of the job satisfaction was trying to do whole. Advocacy went to NPS. Friction went to CES. What remained for CSAT was the middle: did this stage of the journey satisfy you. The three are not three answers to one question. They are three fragments of a question too big for one number.

Satisfied is a weak word, and this needs no statistics

The stronger argument against satisfaction as a headline metric stands entirely on its own: it would survive even if some future study crowned a predictive winner.

Satisfied means nothing went wrong. That is the whole content of the word. A customer who says they are satisfied is telling you the experience cleared the bar of acceptability. They are not telling you they were delighted, which is the state NPS reaches for with the recommendation question. They are not telling you they got what they came for, which is the state CES asks about directly. Satisfaction sits in the flat middle between failure and enthusiasm, and a programme that measures only the flat middle will miss movement at both ends.

"Satisfied only tells you nothing went wrong. It doesn't tell you the customer was delighted, and it doesn't tell you they got what they came for. That is why we ended up with three metrics instead of one."

Amitayu Basu, CEO and Co-founder, Numr

This is the honest case for running more than one metric, and it is not a predictive case. It is a case about detection: different questions detect different states, and no single question detects all three. The practical problem becomes where to point each question.

The three layers of customer satisfaction metrics

The three metrics are not competing answers to one question; they operate at different altitudes.

NPS is the relationship layer. It is asked periodically, not after a specific event, and it measures advocacy: the customer's standing disposition towards the brand as a whole. If you want the fuller treatment, we have written separately on what NPS is and on what counts as a good NPS.

CSAT is the journey layer. It is asked after a stage: onboarding, delivery, renewal, a support episode. It measures whether that stage satisfied the customer, and it localises the signal to a named part of the journey.

CES, which Numr scores as the Net Easy Score, is the process layer. It is asked after a single interaction, and it measures whether the process got in the way. Of the three, effort is the one that tells you what to fix, because a process is something an operations team can change. We cover the mechanics in our CES explainer.

That last point is easily misread as a verdict. Most actionable is not the same as best. Effort tells you what to fix precisely because it sits at the smallest layer, where what is measured is something a team can change. That is a property of the layer, not a superiority of the metric.

The layer model also explains a common programme failure. A programme that points all three metrics at the same layer, three questions after the same support call, say, ends up with scores nobody acts on. The value of running three metrics is coverage of three layers. Pointed at one layer, they are redundancy dressed as rigour.

The definitions, and why the numbers are not the same kind of number

The three formulas look similar, but they produce different kinds of number, and a report that mixes the kinds misleads without anyone intending it to.

CSAT is positive responses divided by total responses, times 100. Commonly that means the share of respondents rating 4 or 5 on a 5-point scale. CSAT is a genuine percentage. A CSAT of 80% means 80% of the people who answered gave a positive rating. It is a percentage of its own respondent base, though, and that is exactly why it cannot be set beside a percentage of a different base: same kind of number, different population underneath.

NPS is the percentage of promoters minus the percentage of detractors. Because it is a difference between two percentages, NPS is a points score on a scale from -100 to +100, and it is not a percentage. Never write it with a % sign. A table that puts "CSAT 80%" beside "NPS 32" is comparing two different kinds of number, and any report that does it should say so in a sentence.

CES, scored as the Net Easy Score, is built by the NET method, a top-group-minus-bottom-group calculation, never an average. Numr's default recommendation is a 7-point scale: the percentage rating 6 to 7 minus the percentage rating 1 to 3, with ratings of 4 and 5 excluded. That default is a recommended starting point rather than a universal rule; the scale is client-driven. Like NPS, NES runs from -100 to +100, and like NPS, it is not a percentage.

A worked example of the CSAT calculation

Suppose a post-delivery survey collects 400 responses on a 5-point scale, all figures illustrative: 180 respondents rate the stage a 5, 140 rate it a 4, 50 rate it a 3, 20 rate it a 2, and 10 rate it a 1.

Positive responses are the 4s and 5s: 180 + 140 = 320.

CSAT = (320 ÷ 400) × 100 = 80%.

The 50 respondents at the midpoint contribute nothing positive to the score, which is right: a 3 on a labelled 5-point scale is the flat middle, and counting it as positive would flatter the stage. And the score is a percentage of respondents, not of customers, a point we return to below.

A worked example of the NPS calculation

Take the same illustrative base of 400 responses, this time to the 0-to-10 recommendation question. Suppose 200 respondents rate 9 or 10 (promoters), 120 rate 7 or 8 (passives), and 80 rate 0 to 6 (detractors). Promoters are 50% of the base and detractors are 20%.

NPS = 50 - 20 = 30.

Not 30%. The passives are simply absent from the subtraction, and the result is a net score rather than a share of anything.

A worked example of the Net Easy Score calculation

On the 7-point default, suppose the same illustrative 400 responses to the effort question: 240 rate 6 or 7, 100 rate 4 or 5, and 60 rate 1 to 3. The 4s and 5s are excluded from the calculation. The top group is 60% of the base and the bottom group is 15%.

NES = 60 - 15 = 45, on the same -100 to +100 scale as NPS, and like NPS it is not a percentage.

Where CSAT belongs now

CSAT is not finished, but its job has changed. CSAT no longer belongs as the headline metric of a programme; it belongs as a sub-question inside a survey, after the recommendation question. For example: "how satisfied were you with your customer support call". In that position its weakness becomes tolerable, because it is no longer being asked to summarise the relationship, only to grade a stage.

Even in that role, the industry has been drifting away from the satisfaction framing towards agree-or-disagree framing, for example "do you agree that customer support was able to solve your problem". An agreement question commits the respondent to something more specific than a feeling. "Satisfied" invites a mood; "was your problem solved" invites a fact.

What keeps CSAT in the toolkit is disarmingly practical. CSAT's real advantage is that a 5-point labelled scale needs no explaining. A customer understands it without instruction, in any market, at any level of survey literacy. In survey design, a question that needs nothing is worth keeping.

Why there is no benchmark in this article

The question a searcher most wants answered next is some variant of "what is a good CSAT score", and the category is happy to answer it with tables. The level of any score is a joint product of category, collection mode and who chose to answer. Two scores are comparable only when all three match, and outside a single programme measured consistently over time, they almost never do. Putting unmatched scores in one table does not make them comparable; it makes the table misleading with confidence.

Numr publishes no CSAT benchmark and holds none, and that is a position rather than a gap in the data. The same applies to CES. The useful question is not whether your score clears somebody else's bar, but whether your own score is moving, measured the same way as last time.

One clause above deserves a final sentence: who chose to answer. Precision comes from the count of responses; representativeness comes from who gave them. A response rate does not tell you the score is right, it tells you who is in the base, and respondents skew towards the involved and the recently affected. We treat that subject properly in our note on survey response rates.

Choosing by layer: the short version

If the predictive contest produced no winner, the selection question is not "which metric is best".

Ask instead: which state do I need to detect, and at which layer? If the question is the standing health of the relationship, that is NPS, asked periodically. If the question is whether a stage of the journey cleared the bar, that is CSAT, asked after the stage, increasingly in agree-or-disagree form. If the question is whether a process is getting in the customer's way, that is CES scored as the Net Easy Score, asked after the interaction.

Three metrics exist because each took a piece of a job satisfaction was trying to do whole, and the right way to deploy them is to give each its piece back. A programme built that way does not need any of its metrics to win a predictive horse race. It needs each of them pointed at the layer it was carved out to measure, scored consistently, and compared only with its own past.

Frequently asked questions

What are customer satisfaction metrics?

Customer satisfaction metrics are the standard survey measures used to quantify how customers experience a company: CSAT for satisfaction with a journey stage, NPS for advocacy at the relationship level, and CES for the effort a process demanded. NPS and CES were each created to cover something satisfaction alone was not capturing.

How is a CSAT score calculated?

CSAT is positive responses divided by total responses, times 100. On the common 5-point scale, positive means a rating of 4 or 5. If 320 of 400 respondents rate a stage 4 or 5, CSAT is 80%. It is a genuine percentage: the share of respondents who answered positively.

Is NPS a percentage?

No. NPS is the percentage of promoters minus the percentage of detractors, which makes it a points score on a scale from -100 to +100. It should never be written with a % sign.

What is a good CSAT score?

There is no defensible universal answer. The level of any score is a joint product of category, collection mode and who chose to answer, so scores are comparable only when all three match, which they almost never do. Numr publishes no CSAT benchmark and holds none, as a position rather than a gap in the data. The useful question is whether your own score is moving, measured the same way as last time.

What is the difference between CSAT, NPS and CES?

They operate at different layers. NPS is the relationship layer, asked periodically, measuring advocacy. CSAT is the journey layer, asked after a stage, measuring whether that stage satisfied the customer. CES, which Numr scores as the Net Easy Score, is the process layer, asked after a single interaction, measuring whether the process got in the way.

Does NPS predict revenue better than satisfaction?

Not demonstrably. NPS was created on the argument that satisfaction did not correlate with anything financially driven, but that is a claim rather than a finding: when a 2007 Journal of Marketing study tested the two against firm revenue growth, neither came out the clear winner. The contest was run and nobody won it.

Should a programme use all three metrics?

Using all three is sensible when each is pointed at its own layer: NPS at the relationship, CSAT at journey stages, CES at individual interactions. Pointing all three at the same layer produces scores nobody acts on. The case for multiple metrics is coverage, not corroboration.

Why is CSAT still worth using if satisfied is such a weak word?

Because its scale is its virtue. A 5-point labelled scale needs no explaining: a customer understands it without instruction, which is not true of a 0-to-10 recommendation scale or an effort question. Used as a sub-question after the recommendation question, and increasingly in agree-or-disagree form, CSAT grades a stage without being asked to carry the whole programme.

Gourab Majumder
Share

See Numr CXM answer your next question.