Survey scale: five points, seven points, or the eleven your hero question already chose

What psychometric research actually shows about 5 and 7-point Likert scales: reliability, labelling, mobile use, and when the choice matters.

· Last updated August 18, 2026
A curled measuring tape on a black background, its graduations and numbers running along the edge.

TL;DR

  • Five-point and seven-point Likert scales both work. On the published evidence, seven is the safer default.
  • A CX scale arrives attached to the hero question: the eleven-point 0 to 10 recommendation question, the seven-point fully anchored effort question, or the five-star rating.
  • Below five points, validity and discriminating power fall away, whatever reliability survives. Two, three and four points are not defensible rating scales.
  • These are three instruments built for different jobs, not three grades of the same instrument.
  • The survey scale of your outcome question shapes what driver analysis can find, before any model runs.
  • Hold one scale length across a survey. Yes-or-no questions are the acknowledged exception.

The three hero questions, and the survey scale each one carries

In customer experience work you rarely choose a survey scale from an open menu. You choose a hero question, and the survey scale comes attached: the eleven-point 0 to 10 recommendation question, the seven-point fully anchored effort question, or the five-star rating. Where the choice is free between a 5-point and a 7-point Likert scale, seven is the safer default and five works nearly as well. Below five, a question stops being a defensible rating scale. The useful comparison is which instrument matches what you intend to do with the data.

Recommendation question

Effort question

Five-star rating

Points

11 (0 to 10)

7

5

Labelling

Words at the two ends only

Every point anchored with a word or icon

None

Best at

Spread: separating customers finely across a long range

Defensible, repeatable measurement of one interaction

Speed and familiarity; highest ease of response

What it costs

Cannot be fully labelled, so interior points are unprotected

Full verbal labelling works against it in regression; every anchor must be translated per fielding language

Least spread; as a reported score, nothing fixes what a star means

Use it for

The relationship outcome question, and driver analysis

Transactional effort measurement, tracked as a Net Easy Score

Quick in-app pulses where completion matters most

A programme can legitimately run all three at once, because the one-length rule is per questionnaire, not per programme: each hero question lives in its own survey.

The recommendation question: 0 to 10

Eleven points is a convention, not a psychometric optimum. The 0 to 10 ask behind the Net Promoter Score carries words at the two ends only, and the length was never the tested part. When Reichheld published the question in Harvard Business Review in December 2003, it had been arrived at empirically: candidate questions were tested against real purchasing and referral behaviour, and ultimately against company growth. That validation attached to the question, not to the number of points it is asked on. Starting at zero gives the scale a true floor for "not at all likely" and a genuine midpoint at 5. The convention is worth keeping because the comparison that matters is a programme against itself over time, and a series only holds if the instrument does not move underneath it.

What eleven points buys is spread. In Preston and Colman's data, discriminating power was best at nine, eleven and 101 points, and respondents rated ten and nine best for expressing what they felt.

Eleven points looks harder to answer than it is. Preston and Colman's respondents rated the ten-point format best of every length tested overall and put it among the easiest to use alongside five and seven: a lifetime of marking things out of ten shows up in the data as fluency with the format.

What it costs is protection. Eleven honest verbal labels do not exist, so the scale runs with endpoint labels by necessity, and everything between the ends is left to the respondent's own reading, without the point-by-point defence a short, fully anchored scale allows. The two end labels still pin the extremes, the only positions where honest words exist at this length, and the respondent constructs the interior on a symmetrical frame.

Note also what the question produces. The Net Promoter Score itself is a score on a scale from -100 to +100, and treating it as a percentage is a common and consequential mistake. When the conversation turns to what number the programme should be producing, that is its own question with its own answer.

The effort question: seven points, every point anchored

The anchors are the point here. As Numr fields the effort question, all seven points carry an anchor, a word or an icon depending on the implementation. The question is scored by the net method, as a Net Easy Score, never as an average: the percentage rating 6 or 7 minus the percentage rating 1 to 3, with 4 and 5 excluded. The banding is a recommended default rather than a universal rule, and the scale is client-driven: other formats are in live use, including a 0 to 10 implementation at a leading digital infrastructure provider, split 9 to 10 minus 0 to 6. On any scale, the Net Easy Score runs from -100 to +100.

Of the three, this one has the strongest evidence behind its measurement, on the criteria of test-retest reliability and response error; on discriminating power the eleven-point scale beats it. One distinction matters. The published evidence is about full verbal labelling. According to Weng, scales with every option labelled with words produced higher test-retest reliability than endpoint-labelled ones. Research from Weijters, Cabooter and Schillewaert (International Journal of Research in Marketing, 2010) found that labelling every category with words reduced extreme responding and reduced errors on reverse-worded items, though it also increased acquiescence, and full verbal labelling is the one design choice shown to anchor a scale against visual and colour cues. Where an icon stands in for a word, no one has tested whether it carries the same protection, and Numr holds no comparison of its own. Full anchoring also carries a practical cost: every anchor word must be translated into every language the survey fields in, and the harder work is establishing anchor equivalence across those languages.

The wording problem that defeats an eleven-point scale has not bitten at seven. No honest words exist for the difference between an 8 and a 9; at seven, a full anchor set can still be written without strain. Nothing establishes seven as the outer limit of full anchoring; the limit simply sits somewhere past it.

The five-star scale

Five stars is the most familiar rating instrument alive; everyone reading this has used one within the week, and it needs no instruction. It is the quickest of the three to answer, reads at a glance, and carries no words, so nothing needs translating when the survey crosses a language. For an in-app pulse where completion matters most, that is a strong case. Even the absence of anchors is not a straightforward defect for analysis: an unanchored scale sits closer to the endpoint-labelled formats that estimate linear relationships well, so for driver work the missing words may be no handicap at all.

What five stars costs is spread and protection: the least spread of the three, and when the number is read as a score, no anchor fixes what any star means. The famous charge against the stars comes from review sites. Hu, Pavlou and Zhang, in a working paper from New York University, studied product reviews on Amazon and found 78, 73 and 72 per cent of book, DVD and video reviews were four stars or better, with just over 90 per cent of products showing a distribution that was neither normal nor unimodal. The decisive part is their experiment: when every student in a class reviewed the same CD, the ratings came out approximately normal, while voluntary Amazon reviews of that same CD were J-shaped. The cause they identify is self-selection. Only favourably disposed people buy, and among buyers, only the strongly opinionated bother to write.

So the J shape lives in who chooses to answer, not in the stars, and an invited CX survey samples differently: you cannot assume a five-star question in your survey will come out J-shaped. What remains is the smaller, unproven suspicion that respondents bring the review-site habit of speaking up only when they feel strongly to a survey that borrowed the same instrument. That is this article's judgement, not a tested finding.

What your survey scale does to driver analysis

The outcome question's survey scale decides how much a driver analysis can find before any model runs. The analysis relates each driver to the outcome, and a coarse outcome variable cannot give back variation it never recorded. Two findings push the same way, a mechanical one about coarseness and an unexplained one about labelling.

The coarseness finding is the mechanical one. Onoshima, Shiina, Ueda and Kubo (Behaviormetrika, 2019) ran a large-scale simulation of what happens to Pearson's r when a continuous variable is cut into categories: the correlation is biased downward, more seriously than earlier work suggested. A larger sample does not repair the bias; it estimates the depressed number more precisely. The paper gives no figure per number of categories, so no threshold is on offer. Preston and Colman's discriminating-power results run the same way, worst at three points and best at nine and eleven. An outcome question on a short scale hands every driver correlation a handicap.

That handicap is at least even-handed: attenuation that lands on all drivers at once, roughly systematically, damages a ranking far less than the magnitudes, and driver analysis is mostly read as a ranking. Coarseness costs the ability to say how much before it costs the ability to say which; the case for spread is strongest where magnitudes matter and not only the order.

The labelling finding is the unexplained one, an association rather than a mechanism. Weijters and colleagues report that for estimating linear relationships, the regressions and correlations driver work is made of, endpoint-labelled scales showed better criterion validity than fully labelled ones. Why full labelling would degrade estimation of a linear relationship is not explained, by the study or elsewhere, and the source does not state what criterion the validity was measured against. Unexplained as it is, it lands awkwardly: the full anchoring that makes the effort question well-measured as a score sits on the wrong side of this finding when that question becomes the left-hand side of a regression.

So the recommendation is plain. If the number will be read as a score, favour full anchoring. If it will be regressed against drivers, favour spread and endpoint labels. The eleven-point recommendation question is built for the second job, which is why it keeps its place as the outcome variable in most driver work. The effort question is built for the first.

Both hero instruments do, on their face, throw spread away at the last step: the recommendation score collapses eleven points into three bands, and the Net Easy Score keeps the top and bottom of seven and discards the middle. The contradiction dissolves on one distinction. The survey scale a question is asked on and the scale it is reported on are different objects. Banding happens at the scoring step, after the response is stored; the raw eleven-point answer is still in the data, and analysis that needs spread runs on the stored responses rather than on the headline score. That is why the eleven points earn their keep even though the score is three bands: a collected response can always be banded later, and a response never collected cannot be recovered by any method.

One survey, one scale length

The house rule is short: hold one survey scale length across the whole questionnaire. Yes-or-no questions are the acknowledged exception: they are routing and fact-checking questions, not ratings, so a rule about rating-scale length does not apply to them.

Dawes found that ten-point data ran a third of a point lower than five-point and seven-point data on a common base, 6.6 against 6.9, significant at p = 0.04, even after rescaling. Five-point and seven-point means convert into each other cleanly, though means are all Dawes established; ten needs a rescale plus an arithmetic adjustment. Weijters and colleagues go further: data obtained with different formats are not comparable, and the format choice must be reported alongside the results. The two are compatible, and the stricter governs: Dawes shows means can be aligned across lengths; Weijters warns the instruments still differ, and nothing establishes the same alignment for the top-box and net proportions CX programmes report on. An alignable mean is no licence to mix formats. A questionnaire that mixes scale lengths puts its own answers on different footings, and a programme that changes length between waves has broken its own series. In tracking work, the movement of the score is the measure: the level on its own is close to meaningless, and movement can only be read against a survey scale that has held still from wave to wave.

The rule itself covers length and nothing else. The rest is ordinary good practice rather than anything Numr prescribes: rating questions take the hero question's length, and an inherited mixed questionnaire is best corrected in a single move at a wave boundary, with the change recorded, so that a future analyst comparing two waves can tell whether the customers changed or the questionnaire did.

Numr has run no experiment on the effect of changing scale length within a live programme. The discipline above rests on the published work, not on a Numr measurement.

Five-point or seven-point Likert scales: what the evidence says

If nothing constrains the choice, seven points is the safer default. Krosnick and Presser, reviewing the question design literature in the Handbook of Survey Research (2010), conclude with a reviewer's caution: "Overall, our review suggests that 7-point scales are probably optimal in many instances." A study by Preston and Colman (Acta Psychologica, 2000) tested every length from two points to eleven, plus a 101-point version, with the same 149 respondents. Reliability was highest at roughly seven to ten points. Asked which formats best let them express what they actually felt, respondents ranked ten, nine and seven top, with five absent from that list.

Five is not wrong. It reads faster, it fits a phone screen without shrinking, and respondents in the same study rated it among the easiest to use. Dawes (International Journal of Market Research, 2008) put the same eight questions to matched samples at five, seven and ten points and rescaled the results to a common base: the five-point and seven-point versions produced identical means, 6.9 out of 10. A 5-point mean translates into a 7-point mean cleanly, so choosing five costs little that a rescale cannot recover. The limit of that finding: Dawes established equal means only; top-box and net proportions, the forms most CX reporting takes, were not tested.

Where a survey scale stops working

The real line for a survey scale is lower down. In Preston and Colman's data, validity and discriminating power fall away together at two, three and four points, and Weng (Educational and Psychological Measurement, 2004), with 1,247 respondents across lengths of three to nine, found that fewer categories cut test-retest reliability hardest. Respondents found those short formats quickest to answer yet rated two and three points extremely poorly for expressing what they felt. The disqualification rests on validity and discriminating power rather than on reliability: reliability is the contested part of the verdict, since not every study agrees, and neither of the other two failures is contested.

The midpoint, longer scales and where the literature disagrees

Keep the neutral midpoint. Krosnick and Presser recommend it, citing O'Muircheartaigh and colleagues (1999): adding midpoints improved both reliability and validity, and Weijters and colleagues found a midpoint also reduces extreme responding, at the cost of some acquiescence. Remove it and genuinely neutral respondents do not become decisive; an even-numbered scale builds the forced pick into the design, and the arbitrary side each chooses enters your data dressed as an opinion.

Ten and eleven points are legitimate survey scale lengths where spread is the goal. They cannot be fully labelled, and Dawes' adjustment applies whenever their results are set beside shorter-scale data.

The colours, the icons and the layout of the points on the screen are a subject of their own, with evidence of their own, and they are covered in the article on how Likert scales are presented.

The survey scale literature does not converge on seven, and the disagreement is worth knowing. Maitland's review in Survey Practice (2009) lays it out: Krosnick and Fabrigar found five to seven points more reliable than shorter or longer scales; Alwin's analysis of more than 300 survey questions found two-point scales the most reliable of all; Saris and Gallhofer argue that up to eleven categories may be optimal. Alwin's result sits oddly beside the floor above until the criteria are separated: a two-point scale can repeat its answer perfectly and still be unable to tell a delighted customer from a merely satisfied one, which is the discriminating-power failure Preston and Colman record. Maitland's own recommendation is to pretest with your population rather than pick a number from a table. Length alone does not settle the matter. Start from the question and the use of the data; in most CX programmes, that means accepting the scale your hero question already chose.

Frequently asked questions

Is a 5-point or a 7-point Likert scale better?

Both work. Seven is the safer default: Krosnick and Presser call seven-point scales "probably optimal in many instances", and respondents in Preston and Colman's study ranked seven among the best formats for expressing what they felt. Five is quicker and fits small screens, and Dawes showed five-point and seven-point results rescale into each other with identical means.

Why is the recommendation question asked on a 0 to 10 scale?

By convention, not because eleven points is a proven optimum. The 0 to 10 format is the one the Net Promoter Score was introduced with, and it stuck because recommendation programmes are compared against their own history; starting at zero gives the scale a true floor and a genuine midpoint at 5. The length also buys spread: discriminating power in Preston and Colman's data was best at nine and eleven points.

Should a rating scale have a neutral midpoint?

Yes. Krosnick and Presser recommend keeping it, citing evidence that midpoints improve reliability and validity. Without one, genuinely neutral respondents are forced to pick a side arbitrarily.

Can I compare results collected on different survey scale lengths?

Not directly. Dawes found five- and seven-point data rescale into each other with equal means, though means are all he tested, while ten-point data ran measurably lower even after rescaling and needs an arithmetic adjustment. Weijters and colleagues state that data from different formats are not comparable and that the format used must be reported.

Do five-star ratings always come out J-shaped?

No. The J-shaped distributions Hu, Pavlou and Zhang found on Amazon came from self-selection, who buys and who bothers to review, not from the stars: when a whole class rated the same CD, the distribution was approximately normal. An invited survey samples differently from a review site.

How many points is too few?

Anything below five. At two, three and four points, validity and discriminating power sit at their worst, whatever reliability survives, and respondents rate those formats poorly for expressing what they feel. Yes-or-no questions are fine; they sit outside the rating-scale rules as routing questions.

What is a Net Easy Score?

The net method applied to the effort question: the percentage of respondents rating the interaction easy minus the percentage rating it hard, giving a score from -100 to +100. On Numr's seven-point effort question the recommended default banding is 6 to 7 minus 1 to 3, with 4 and 5 excluded; the banding and the scale are client-driven.

Will a bigger sample make up for a coarse survey scale?

No. Onoshima and colleagues showed that categorisation biases correlations downward and that a larger sample only estimates the depressed value more precisely.

Share

See Numr CXM answer your next question.