What Is Customer Effort Score? And What Your CES Number Cannot See

Customer Effort Score (CES) measures how easy it is to deal with you. See what CES is, the Net Easy Score formula, what a good score is, and how to act on it.

Amitayu Basu CEO and Co-founder
17 min read
Abstract Numr artwork: a row of analogue gauges, two lit.

The short answer

Customer Effort Score measures how hard a customer had to work to get something done: resolve a problem, complete a purchase, get an answer. Numr scores it as a net figure rather than an average, subtracting the percentage who found the interaction difficult from the percentage who found it easy. We call that the Net Easy Score.

The choice of method is not cosmetic. An average of effort ratings cannot reliably tell you whether this quarter was better than last, because the answer depends on what the numbers on the scale are assumed to mean. A net score does not have that problem, and it also refuses to let the customers who struggled disappear into the middle.

Everyone calculates it the same way

Read the top-ranking guides on customer effort and you get the same instruction. IBM says the score "is simply an average of responses that is calculated by the total sum of each response divided by how many responses there were." Zendesk gives you an average for numbered scales and reserves the net method for emoji.

Add the ratings, divide by responses, report the result. That produces a number like 5.5 out of 7, and it sounds respectable.

What effort was supposed to tell you

Customer Effort Score comes from research published in Harvard Business Review in July 2010 by Matthew Dixon, Karen Freeman and Nicholas Toman, then at CEB. In HBR's own summary of that paper, the authors "introduce the Customer Effort Score and show that it is a better predictor of loyalty than customer satisfaction measures or the Net Promoter Score."

That is the claim the whole metric rests on. It is a claim about a relationship between effort and loyalty, which means the interesting population is not the average customer. It is the group at the difficult end, however small.

An average is a poor instrument for finding that group, for two separate reasons. The first is about the scale. The second is about what averaging does to any two groups.

Problem one: the numbers on the scale are labels, not amounts

When a customer picks the sixth position on a seven-point effort scale, that 6 is not a measured quantity of effort. It is a position in an ordered list, and we chose what to call it. We could have called it 5, or 60.

Averaging only works on amounts, because adding requires the gap between each pair of values to be the same size, and on an ordered list of positions any labelling that keeps the order is as valid as any other. That is not a philosophical point. It changes the answer.

Here are the 2 quarters below, the same 2,000 answers, under 3 labellings that all preserve the order exactly:

Buttons labelled

Q1 average

Q2 average

Which quarter was better?

1 to 7

5.58

5.41

Q1. It got worse.

0 to 6

4.58

4.41

Q1. It got worse.

1, 2, 3, 4, 5, 6, 20

10.13

11.26

Q2. It got better.

The third row is not a trick. It keeps every customer in the same rank order and only says that the top of the scale sits further from the rest than the other steps do, which for effort is a perfectly reasonable thing to think. And it reverses the verdict. The identical customers, the identical answers, and the average now reports the opposite conclusion about whether the journey improved.

You do not have to believe that third labelling is the right one. You only have to accept that nobody has shown it is wrong. The average needs that question settled before it can rank two quarters against each other. The count does not.

Count instead, and this does not happen:

Q1   easy 60%  difficult 10%   NES = +50
Q2   easy 60%  difficult 20%   NES = +40

The net score is +50 and +40 under every one of those labellings, because band membership does not care what is printed on the buttons. So the average cannot tell you which quarter was better without an untested assumption about spacing, and the count can.

Problem two: the happy customers pay for the unhappy ones

This one has nothing to do with scales. It is what averaging does to any two groups: it lets one group's gains pay for another group's losses. The consequence is that you cannot work the number backwards. 5.5 out of 7 is consistent with 5% of your customers in difficulty and equally consistent with 20%, and here are the two distributions that prove it:

700 rated 6, 250 rated 5,  50 rated 2      average 5.55    difficult  5%
450 rated 7, 350 rated 5, 200 rated 3      average 5.50    difficult 20%

Same headline, four times the problem. Nothing in the number tells you which one you have.

Here is what that looks like across 2 quarters of the same support journey, 7-point scale, 1,000 responses each. Numbers you can check.

Rating

Q1

Q2

7 very easy

350

450

6

250

150

5

200

130

4

100

70

3

50

80

2

30

70

1 very difficult

20

50

Total

1,000

1,000

The averages:

Q1   (7×350 + 6×250 + 5×200 + 4×100 + 3×50 + 2×30 + 1×20) / 1000  =  5.58
Q2   (7×450 + 6×150 + 5×130 + 4×70  + 3×80 + 2×70 + 1×50) / 1000  =  5.41

The net scores over the same 2 quarters:

Q1   easy (6-7) 60%   difficult (1-3) 10%      NES = 60 - 10 = +50
Q2   easy (6-7) 60%   difficult (1-3) 20%      NES = 60 - 20 = +40

So the average fell 0.17 of a point, from 5.58 to 5.41, while the number of customers in real difficulty went from 100 to 200. It doubled.

0.17 against a doubling. That is the distance between what the average reports and what happened, and it is not a quirk of these particular numbers. It follows from what an average does: it lets a growing group of delighted customers pay for a growing group of struggling ones and reports the settlement. Commercially those two groups do not cancel at all. The delighted ones were already staying. The miserable ones are the ones who leave.

What the net method cannot see either

The net method has its own blind spot, and it is worth stating plainly because most of the people selling it will not.

Take Q2 above and move its 80 threes down to a 1. That is the worst deterioration available on the scale, customers who found it difficult now finding it as hard as it gets. Nobody enters or leaves the difficult band. They just sink inside it.

Q2

Q2 after the slide

Rating 3

80

0

Rating 2

70

70

Rating 1

50

130

Average

5.41

5.25

Net Easy Score

+40

+40

% difficult (1-3)

20%

20%

% at the very bottom (1)

5%

13%

The average catches it. The Net Easy Score does not move at all, because a 3 and a 1 sit in the same band and banding cannot see inside itself. Worse, the difficult percentage does not move either. Those 200 customers were already counted.

So neither headline is the truth, and neither is the count behind it. The average dilutes the difficult group into the middle. The net method surfaces the group and then stops paying attention to what is happening inside it.

Push it once more and the point generalises. Move the same 80 customers to a 2 rather than a 1:

Q2 baseline          average 5.41   NES +40   difficult 20%   at the bottom 5%
80 slide from 3 to 1 average 5.25   NES +40   difficult 20%   at the bottom 13%
80 slide from 3 to 2 average 5.33   NES +40   difficult 20%   at the bottom 5%

In that third row the net score, the difficult percentage and the count at the very bottom all sit exactly where they were, and only the average moves. Any single summary is lossy. That is not a reason to pick none of them.

It is a reason to state the rule properly. Manage the percentage of customers who found it difficult, and how far down the scale that group sits. A difficult band can get worse without getting bigger, and its size is what the guides above report.

There is an irony in that worth owning. Tracking where the difficult group sits means taking an average, inside the band. The average is not a bad calculation. It is the wrong headline and the right microscope, and the mistake the whole category makes is running it across everybody instead of across the people in trouble.

It inherits the labelling problem too, so read it the way you read any trend: hold the labels fixed and watch whether the position moves. The direction is trustworthy even where the magnitude is not, which is the same reason the level of any score matters less than its movement.

The net method's real virtue, then, is not that it sees everything. It is that you cannot produce it without computing the difficult percentage first, which puts that number in front of whoever reads the report. An average across everybody can be calculated, presented and discussed for 4 quarters without anyone in the room ever learning it.

The Net Easy Score

Numr scores customer effort using the net method, which we call the Net Easy Score, or NES.

On the recommended 7-point scale, where 7 means the interaction was very easy and 1 means it was very difficult:

NES  =  % who rated 6 or 7   -   % who rated 1, 2 or 3

The middle scores, 4 and 5, are excluded. Worked on Q1 above:

NES  =  60%  -  10%  =  +50

Like NPS it runs from -100 to +100.

The 7-point scale is a recommendation rather than a rule. In practice the scale is usually driven by the client, and other scales are in live use. On a 0 to 10 scale:

NES  =  % who rated 9 or 10   -   % who rated 0 to 6

with 7 and 8 excluded as the middle, which makes it structurally identical to NPS and easy to read alongside one.

Banding makes an assumption of its own, and it is only fair to name it. It treats 1, 2 and 3 as interchangeable for the purpose of deciding what to do, and treats 4 and 5 as carrying nothing worth reporting. That is not free either, and the counterexample above is exactly where it bites. Where you draw the line is still a choice, and there is no deep reason the cut falls between 3 and 4 rather than somewhere else.

The difference is that it is a choice you can see and argue about. Averaging's assumption is not a grouping policy you can put on a slide. It is a claim that the numbers on your buttons behave like quantities, and the verdict-flip above is what happens when that claim is wrong.

The question to ask

The wording Numr uses is:

How easy was it to [do the thing]?

The thing is the specific task the customer has just attempted, named concretely: find the right page on the website, open an account at the branch, get your problem solved by customer support. The task belongs in the question. A generic version like "how easy is it to do business with us" fails, because the customer does not know which part of the business you mean, and neither will you when you read the answer.

Very difficult sits on the left of the scale, very easy on the right, so the score rises as effort falls.

Two rules about what you may combine with it. A new effort question does not absorb your existing NPS or CSAT responses. Those are different questions, and nothing turns an answer to one into an answer to the other.

Older effort responses collected on a different scale can be pulled into the new series. What travels is band membership, not the rating: a customer who sat in the easy band on the old scale sits in the easy band on the new one, under a mapping agreed once and written down. That works because the question was the same and only the ruler changed, and because a net score never needed the rating itself. The old numbers are not rescaled into new numbers, which would run straight into the labelling problem above.

One caveat, and it is the same one we apply to benchmarks. People use a 7-point scale differently from an 11-point one even when the bands are mapped honestly, so mark the changeover on the chart and read the join the way you would read any method change. Re-sorting is the best available continuation. It is not a clean one.

Timing follows from the same logic. Effort is a property of a specific interaction, so ask immediately after the one you named: the call ends, the claim is filed, the return is processed, the account is opened. Ask about the relationship in general instead and you are no longer measuring effort at all. You are measuring how the customer feels about you overall, which is a different question with a different metric attached to it.

If you are switching from an average

There is an obvious objection to all of this. If comparability depends on holding the method constant, then changing method breaks your trend, and you have just been told to change method.

It need not, provided you still hold the individual responses. Banding is a function of raw responses rather than of the summary, so if you have kept them, historical quarters can be recomputed under the new banding and the series is continuous. Recompute back as far as the raw data goes, publish the net series from the start, and keep the average alongside it for a period if people are attached to it.

Two conditions on that. The recomputed trend only means something if the question wording was constant across those quarters. A change of scale is survivable, because old responses can be re-sorted into the new bands as described above, but a change of question is not, and it is worth checking which of the two you had before anyone presents the series.

And if your programme kept only exported summaries, or the retention window has closed, the history cannot be rebuilt. The honest move then is to start the net series fresh and say so.

One gap worth naming: if you are on a 5-point scale, which Zendesk among others recommends, the banding above does not apply as written and the split has to be agreed before anything is re-sorted.

So what is a good CES?

There is no such thing, and Numr does not publish one.

That is not a gap in our data. It is the same position we hold on every score. The level of any score is a joint product of the scale you used, the journey you measured, the moment you asked and who chose to answer. Change any one and the number moves without anything about your service changing. Two effort scores are comparable only when all four match, which between two companies they essentially never do.

So a good NES is not a number anyone can hand you. The useful question is whether your own number is going up, measured the same way as last time. Anyone publishing a CES benchmark table is either averaging unlike programmes together or quoting panel research, and in both cases you are being handed a figure that will not survive contact with your own data.

The practical consequence is that whatever scale and banding you agree on, you hold constant afterwards. A method that is consistent and slightly imperfect beats a method that is theoretically better and changes between quarters, because only the first lets you read a trend.

When to use effort rather than NPS or CSAT

The three metrics are not competing answers to one question. They sit at different levels of the relationship, and a programme that points all three at the same level ends up with a drawer full of scores nobody acts on.

NPS belongs at the relationship layer, asked periodically, answering whether the customer is still with you in spirit. CSAT belongs at the journey layer, asked after a stage, answering whether that stage satisfied them. Effort belongs at the process layer, asked after a single interaction, answering whether your process got in the way. Effort is the one that tells you what to fix, because a process is a thing an operations team can change on Monday.

The full comparison, including which to run at which stage and how to read them together, is in NPS vs CSAT vs CES: which CX metric should you use.

From a score to something you can fix

A Net Easy Score tells you the size of the problem, not its location. Two things get you to the location.

The first is the open text behind the difficult ratings. Customers who rated 1 to 3 have usually said why, and those verbatims are pre-filtered to people who actually hit something, so they are more concentrated than feedback collected from everyone.

The second is driver analysis run against the effort question, which tells you which stages of the journey are moving the score rather than which stages generate the most complaints.

Then the loop closes in the ordinary way. Fix the stage, hold the measurement method still, and watch the 2 numbers from earlier: the size of the difficult group, and where it sits.

In practice the share at the very bottom of the scale is the easiest version of the second one to put on a dashboard, and it is a good proxy most of the time. When either number moves and you want to know what actually happened, look at the shape of the difficult band itself rather than at any summary of it.

---

Frequently asked questions

How is Customer Effort Score calculated?

Numr calculates it as a net figure. On a seven-point scale, take the percentage of customers who rated the interaction 6 or 7 and subtract the percentage who rated it 1, 2 or 3. Scores of 4 and 5 are excluded, and the result runs from -100 to +100. Most published guides tell you to average the ratings instead, which produces a number that can fall by fractions of a point while the population in difficulty doubles.

What is a good Customer Effort Score?

There is no universal good score and Numr does not publish a benchmark. A score's level depends on the scale used, the journey measured, the moment it was asked and who chose to answer, so two companies' effort scores are almost never comparable. Ask instead whether your own score is improving, measured the same way each time.

What scale should a CES survey use?

Numr recommends 7 points, where 7 is very easy and 1 is very difficult, with very difficult on the left of the scale and very easy on the right. It is a recommendation rather than a rule, and the scale is often set by the client. A 0 to 10 scale is also in live use, banded as 9 to 10 easy against 0 to 6 difficult. A scale change is recoverable by re-sorting old responses into the new bands, though the join should be marked on the chart. A change of question wording is not recoverable that way.

Is CES better than NPS?

Neither is better. CES is asked about a single interaction and tells you whether a process got in the way. NPS is asked about the relationship and measures advocacy. The 2010 HBR research that introduced effort argued it predicts loyalty better than satisfaction or NPS, but in practice the two run together: effort at the process level, where something specific can be fixed, and NPS at the relationship level. The full comparison is in NPS vs CSAT vs CES: which CX metric should you use.

Why not just use the average?

Because an average assumes the distance between every pair of adjacent ratings is equal, which has never been established for effort, and because it can be reported without anyone computing how many customers found the process difficult. That percentage is the number the metric exists to surface.

Amitayu Basu CEO and Co-founder

25 years in customer experience, helping global brands listen. Numr is what he built when listening stopped being the hard part.

Share

See Numr CXM answer your next question.