Rating scale presentation: how the format changes the number

A Likert-style rating scale measures opinions or attitudes using ordered responses (for example: strongly disagree → strongly agree).

· Last updated August 18, 2026
A row of mixing-console faders, each set at a different position along its track.

A rating scale asks people to place a judgement on a fixed set of ordered points: how strongly they agree, how satisfied they are, how likely they are to recommend. Its canonical form is the Likert scale, the strongly-disagree-to-strongly-agree format every survey respondent has met. This page is about how to present a rating scale once you have written it: the control the respondent touches, the layout, the labels, the icons, the colour, the device. Each has been studied. Every one of them can move the number you collect, and none has anything to do with your service.

TL;DR

  • Use buttons or click-anywhere zones for a rating scale; sliders produce more skipped items and lower means.
  • Run a horizontal row to seven points. At ten or eleven on a phone, use a vertical stack or a tap-based radial arc, whichever fits the questionnaire.
  • Label every point up to seven. At ten or eleven, honest labels run out: label the endpoints and weigh the remaining choices more carefully.
  • Smiley faces are safe and well liked, and they help respondents who read least fluently. Avoid animated faces.
  • Use no colour or a monotonic gradient. Scoring-band colour turns an eleven-point question into a three-point one.
  • Record every choice and treat any later change as a break in the series, not a movement in the service.

The choice

Use

Why

Control

Buttons or click-anywhere zones

Least non-response; nothing to anchor to

Layout, five to seven

Horizontal row

Fits any screen, whole range visible

Layout, ten or eleven on a phone

Vertical stack or tap-based radial arc

A row runs out of width; both fit a thumb

Labels, to seven points

A verbal label on every point

Cancels visual-cue effects entirely

Labels, ten or eleven

Endpoint labels with numbers

Eleven honest labels do not exist

Icons

Static faces if wanted, never animated

Well liked, no shift, faster for low-literacy readers

Colour

None, or a monotonic gradient

Direction without boundaries; bands rewrite the question

What a rating scale is, and where the Likert scale fits

The Likert scale takes its name from Rensis Likert, the American social psychologist who introduced it in a 1932 paper: respondents reacted to a series of statements on one ordered response set, combined into a single score.

Strictly, a Likert scale is the battery: several statements about one underlying attitude, each answered on the same set of points, summed into one measure. A single statement with its response options is a Likert item. Most of what CX programmes call a Likert scale, a lone CSAT-style satisfaction question after a support call, is in fact a Likert item. A summed battery smooths out any single item's quirks; a lone item takes them at full strength. A rating scale example, invented for illustration: "The support agent resolved my issue", answered on Strongly disagree; Disagree; Neither agree nor disagree; Agree; Strongly agree.

The response options are the points, ordered and nothing more. The data a Likert item produces are ordinal: a 4 sits above a 3, but nothing guarantees the distance from agree to strongly agree equals the distance from neutral to agree. Taking a mean quietly assumes those distances are equal. For a summed multi-item scale the assumption is usually tolerable; for a single item it is exposed, and the safer habit is reporting the full distribution alongside any average.

One behavioural quirk comes with the format: acquiescence. People agree more readily than they disagree, whatever the statement says. The standard defence is to reverse-phrase some items, so that agreement sometimes indicates a negative attitude, which also catches respondents ticking one column all the way down. The defence has a price: reversed items are harder to read, and a respondent who misses the negation produces an error the analysis cannot see.

The other standing question is the midpoint: whether to offer a neutral option, and what respondents do with one. That question is tangled up with scale length, and both belong to a separate article on choosing the number of points; nothing here settles them. Cut it down far enough, though, and it stops being a Likert scale and becomes a yes-no question.

The control: buttons, click-anywhere and sliders

Use buttons. The most direct comparison available tested rating scale items as radio buttons, sliders and click-anywhere areas across desktop, tablet and phone, at five, seven and eleven points (Toepoel and Funke, 2018). Buttons produced the least item non-response, click-anywhere performed as well, and sliders lost twice: considerably more skipped items, and lower means.

The mechanism, in the common implementation, is the default position. A slider has to start somewhere, and wherever it starts is a suggestion: a respondent who does not care much leaves the handle where it is or abandons the item, and the measurement is dragged toward the start position or out of the data. A button has no default, so there is nothing to anchor to. A drag also demands motor precision a tap does not, a real cost on a phone held in one hand. A slider built to dodge all this, no handle drawn until first touch, snapping to discrete points, is a different object from the one usually fielded; the measured result above concerns the usual one.

The case for sliders is made in every design review, so hear it fairly. They look modern. Run as a continuous track, they permit fine-grained response. Stakeholders like them in demos. None of that is imaginary; it is simply priced. The price is more missing answers, a mean pulled toward wherever the handle rests, and scores not comparable with a button-based series. The fine grain is worth less than it appears: a respondent holds an attitude to the nearest category, not the nearest hundredth, so continuous precision is mostly noise. Stated that way, few programmes buy it.

Click-anywhere is the quiet compromise. It behaves like a button with a larger target: the respondent taps anywhere inside a zone rather than a small circle, which suits thumbs, and carried no penalty in the comparison. Where tap accuracy is a worry, it is the same measurement with more forgiving ergonomics.

Layout: row, stack, grid, radial, stars

A horizontal row of tappable points is the default rating scale layout for a reason, and it fails at exactly one predictable place: a long scale on a narrow screen. The alternatives handle that failure or serve many items at once, each with a bill.

The horizontal row

A row is the 1 to 5 rating scale's home ground. The whole range fits on any screen, and seven points still do. At ten or eleven points on a phone the arithmetic turns against it: eleven targets across a narrow display leave each one small, and mis-taps rise. The row is not wrong at eleven points; it is cramped, and cramped is where errors live.

The vertical stack

A stack survives where a row cannot. Each point gets a full line at any length, and the format costs nothing in familiarity, since everyone scrolls vertical lists. It solves the same narrow-screen problem as the radial arc below. The costs are scrolling past unseen points and, when stacked items run in sequence, straightlining down a column. One note covers the whole geometry question: row against stack against arc rests on ergonomics rather than experiment.

The grid or matrix

A grid earns its place when several items share one construct, with statements down the side, responses across the top, and one screen doing the work of six. It is also the format respondents most dislike. In the largest comparison available, 7,096 Dutch public-sector employees answered one of eight designs; radio-button grids came last on questionnaire experience, 4.04 out of 5 (Toepoel, Vermeeren and Metin, 2019). The same study produced a result nobody has explained: tile grids, each cell a large tappable tile, yielded the highest satisfaction scores of all eight formats, 3.60 to 3.64. One caveat covers every number from this study, here and below: means only, no significance tests, intervals or corrections reported, and with eight formats something has to finish last.

Radial and circular scales

A radial scale exists for one reason: width. Bending the points into an arc or a ring fits eleven of them inside the sweep of a thumb, where a straight row on the same phone cannot. At ten or eleven points on a phone, then, use a vertical stack or a tap-based arc, and choose between them on fit: a stack suits a questionnaire that already scrolls, an arc suits a single item on one screen.

One rule decides whether a radial scale is defensible: it must be a tap, not a drag. A ring the respondent drags around is a curved slider, and it inherits everything the previous section said about sliders. A ring of discrete tappable points is a Likert row bent into an arc, and carries no such baggage.

Against a straight row, radial costs familiarity: respondents have seen thousands of rows and few arcs. At five or seven points that cost buys nothing.

Stars and hearts

Stars arrive pre-loaded. Every review site on the internet has trained respondents to read five stars as a specific kind of judgement: useful if a five-point review-style rating is what you want, a liability otherwise. Hearts are a caution. In the same study they produced the lowest satisfaction score of the eight designs, 3.33 against 3.53 for smileys, with no significance test reported (Toepoel, Vermeeren and Metin, 2019). One study, one low number: not proof that hearts depress scores, but reason enough not to assume a decorative choice is a neutral one.

Labels: the choice everything else depends on

Label every point of a rating scale if you can. The evidence here is unambiguous. In a series of web experiments on scale shading, respondents shifted their answers when the two ends of a scale carried different hues, and the shift disappeared entirely when every point carried a verbal label (Tourangeau, Couper and Conrad, 2007). The mechanism the authors describe is simple: respondents treat incidental visual features as meaning when the words are missing, and stop doing so when the words are there, which is the protection that full labelling buys.

At ten or eleven points, you cannot label every point. The barrier is physical first: eleven legible labels will not fit on a phone screen, and in Numr's programmes the great majority of surveys are answered on one, an observation across the book rather than a counted result. It is semantic second, and this is the harder wall: no language offers ten or eleven distinct, ordered, commonly understood degrees of satisfaction, so no honest wording separates a 9 from an 8. Try to write eleven graded expressions and you produce filler by point four; filler labels are worse than none, because the respondent tries to take them seriously and the instrument starts measuring their reading of your wording. The wall also sits in a different place in every language, a further problem for any programme fielding one instrument in several languages at once.

That leaves endpoint labels with numbers, or numbers alone, and either way the scale is permanently without the protection Tourangeau and colleagues identified. That is not an argument against long scales; whether eleven points buy you anything belongs to the scale-length article. It is an observation about weight: on a 1 to 5 rating scale, presentation choices sit under the cover of full labelling; at eleven points there is no cover, and every remaining choice of icon, colour and device carries more load.

Icons and the smiley face rating scale

The common worry about a smiley face rating scale, that it makes people answer more warmly, is not supported. An eye-tracking study found response distributions did not significantly differ between face and no-face versions of the same satisfaction rating scale (Stange, Barry, Smyth and Olson, 2016). In the eight-format study, smileys sat with plain radio buttons on the substantive score, 3.53 against 3.44 to 3.51, and were the best-liked format at 4.42 out of 5 (Toepoel, Vermeeren and Metin, 2019).

Two costs are real. First, faces change where people look: respondents shown faces spent about 1.4 seconds less fixating on the text labels on average, a saving of about 0.94 seconds for high-literacy respondents and 1.90 for low-literacy respondents, roughly twice the benefit for those who read least fluently, a result measured so far only in populations with a narrow literacy spread (Stange et al., 2016). Before filing the fixation saving as a defect: a low-literacy respondent who gets through the item faster and produces the same distribution has had the reading done for them, an accessibility gain. If you believe the labels carry nuance the faces lack, the same numbers describe a loss. Both readings stay open: the study reports a null, not equivalence, and an unmoved aggregate can hide offsetting individual movement. Second, animation costs time: animated face designs took about 14 seconds per item against 9 for a radio-button control across 611 respondents, with nothing to show for it (Emde and Fuchs, 2012).

Then an idea worth testing: a hypothesis with evidence attached, not a finding. A graduated face set is doing the job of a verbal label: the frown deepens toward zero, the mouth flattens at the middle, the smile widens toward ten. Not one smiley repeated eleven times: a face that changes at every point, so each point carries its own meaning rather than being inferred from position alone. Carrying meaning at every point is what a verbal label does, and the previous section established that verbal labels are unavailable at this length. Graduated icons are a candidate substitute for a protection otherwise impossible on a long scale.

The supporting evidence is real, and it is American. In US-only work, respondents rated individual smiley designs on a 0 to 100 scale, and carefully designed faces spread fairly evenly along it, carrying ordered, roughly interval meaning (Sedley, Yang and Hutchinson, 2017). The same work found the scaling improved when endpoint verbal labels were added; there, faces and words behaved as complements. A follow-up fielded the faces in six countries (Sedley, Yang and Paxton, 2020); its published abstract reports the method, not the outcome, so whether the result travels is not yet on record. And the 2017 work shows ordered numeric meaning, not movement in a score.

One test would settle it: a graduated-face scale against a fully labelled one on the same item. Numr has no data of its own on this either: clients choose one presentation and keep it, so Numr has never run smiley against non-smiley on the same programme and cannot say whether faces move a score. The practical decision does not wait on the hypothesis: use faces if you want them, plain or graduated, skip animation, and record the choice.

Colour: none, gradient, or bands

Colour on a rating scale moves answers less than the argument around it would suggest. The foundational result already appeared in the labels section: different hues at the two ends of a scale shifted responses toward the high end, the effect was smaller than the shift produced by negative number labels, a scale running minus five to plus five rather than nought to ten, and it vanished under full verbal labelling (Tourangeau, Couper and Conrad, 2007). In practitioner testing, not peer-reviewed and weighed accordingly, an NPS item run with no colour, a gradient, and three-colour banding across 229 respondents produced detractor shares of 32%, 26% and 24%, patterns the author described as similar (Sauro, MeasuringU, 2019). Those figures measure the share answering 0 to 6, not the mean, and 32 against 24 is a quarter of the detractor base; the similar-patterns reading is the source's own.

Of the three options, no colour is the uncontested baseline. A monotonic gradient, one hue deepening from end to end, carries direction and nothing else: which end is good, no boundary anywhere. It is Numr's default treatment. A constant gradient also produces no ongoing shift wave after wave, because there is nothing for it to shift from: the 2007 effect was measured against an uncoloured control. It can still sit at a constant offset against an uncoloured scale, which matters when comparing programmes, though the trend is untouched.

Banding is different in kind. Some programmes colour the points by scoring band: 9 and 10 green, 7 and 8 amber, 0 to 6 red, mirroring the NPS cut-points. The respondent has no idea what NPS is and does not know promoters or detractors exist. They have no reason on earth to think a 6 is red and a 7 is amber; that boundary lives only in the head of whoever knows the scoring rule. Painting it onto the scale collapses eleven points into three and fixes where the cuts fall. What is presented as a 0 to 10 question is operationally a red, amber and green question with numbers written on it. The essence of the question has changed and the respondent was not told.

Three things keep this claim honest. First, it is a logical claim about what the instrument is, not a claim about movement; it needs no study. The no-drift defence for the gradient covers banding equally: a programme banded from wave one drifts no more than a gradient one. The objection was never drift; it is that the instrument is not what it appears to be, on the first wave and every wave after. Second, Numr's own observation, direction only: some programmes have asked for banding, Numr has fielded it, and the score went up. No magnitude, and the comparison was uncoloured against banded, not gradient against banded, so it cannot separate colour as such from visible cut-points. Third, the experiment that would settle it, gradient against banded on one programme over one period, is unrun, by Numr or anyone else. The decision does not wait on it: run no colour or a monotonic gradient, both defensible, and keep your pick. If a client asks for bands, tell them what changes: from that wave the question is a three-category one and its series starts fresh.

The device is part of the instrument

A phone changes the score. In the study that settled the button question, desktop respondents produced lower scores than mobile respondents on identical rating scale items (Toepoel and Funke, 2018). In Numr's own programmes the great majority of surveys are answered on a phone: an observation across the book, not a counted result, and the reason the finding matters.

The consequence is conditional. If a programme's device mix changes, for whatever reason, the score can move with it while every design choice stays frozen and the service stays as it was. Why a mix changes is not something this page will guess at; that it can is enough. Log the device mix every wave, report the split next to the score, and when the number moves, check the mix first. A shift that tracks the mix is composition, not sentiment.

The closing rule

Every rating scale presentation choice on this page is a method choice. Swap buttons for a slider, trade plain points for hearts, recolour a gradient into bands, or simply let the device mix shift, and the number can move with nothing changing in the service. What cannot be defended is changing the presentation mid-programme and reading the movement as a change in customer experience.

The recommendations above are defensible; holding your choice constant matters more than which one you made. So choose deliberately and write down exactly what you chose: the control, the layout, the labels, the icons, the colour. Treat any later change as a break in the series. Annotate the break on every chart that crosses it, and do not report the step as a result. A programme that does that can use almost any presentation here. A programme that does not will eventually celebrate a redesign.

Frequently asked questions

What is a Likert scale?

A Likert scale is a rating scale that collects responses to statements on a fixed set of ordered points, most familiarly strongly disagree to strongly agree. Strictly, the scale is a battery of such items combined into one score; a single question with its options is a Likert item.

Should I use a slider or buttons for a rating scale?

Buttons. In direct comparison, sliders produced considerably more item non-response and lower means; in the common implementation the handle's starting position acts as an anchor. Click-anywhere areas performed as well as buttons, so either is safe.

When does a radial scale beat a horizontal row?

At ten or eleven points on a phone, where a straight row runs out of width. An arc fits the full range within a thumb's reach; so does a vertical stack, at the cost of scrolling. It must be discrete taps, not a drag, or it becomes a curved slider.

Do smiley faces bias survey results?

The evidence points to no: response distributions did not significantly differ with and without faces, and a smiley face rating scale matched plain radio buttons across 7,096 respondents. Faces reduce time spent reading labels, most for low-literacy respondents. Animated faces, though, roughly double response time for no benefit.

Should a rating scale be coloured?

No colour and a monotonic gradient are both defensible; measured colour effects on rating scales for surveys have been modest. Colouring by scoring band is the option to treat with care, because it draws boundaries the respondent cannot expect and quietly converts an eleven-point question into a three-category one.

Can I change my rating scale design mid-programme?

You can, but treat the change as a break in the series. Record the old and new presentation, annotate every chart that crosses the change, and never read the step as a movement in customer experience. If it matters enough, run old and new in parallel before switching.

Share

See Numr CXM answer your next question.