Survey response rate: the funnel behind the number
TL;DR
- A survey response rate is the share of people invited to a survey who answer it. The catch is the word invited: the rate is the end of a funnel, and most published rates never say which stage they divide by.
- Answered over delivery-confirmed and answered over actual send-outs are two readings of the same programme; the send-out version is never higher, and lower whenever any message fails to confirm. Quote a rate with its denominator or not at all.
- Records are excluded before anything is sent: invalid contacts, frequency caps, opt-outs, dispatch failures. The base measured is already not the base the client's system handed over.
- The two biggest drivers of response are involvement with the category and the moment of the ask. Neither predicts alone, and a weak reading on either caps the rate: high involvement with no moment still gets very little. Speed, language, channel and format are levers under those two, in that order.
- A low rate is a symptom with two diseases: losses before delivery are a list and authentication problem, losses after delivery a relevance problem, and the single percentage cannot tell you which you have.
- Response rate and completion rate are different metrics: replies over invitations versus finished surveys over started ones. Conflating them flatters the surveys that are quietly bleeding respondents.
The number is the end of a funnel
A survey response rate is replies over invitations, and both of those words are choices rather than facts. What counts as a reply and what counts as an invitation are two lines drawn through a funnel, and the same programme yields several different, defensible rates depending on where the lines are drawn. That is a property of the arithmetic, not a scandal, and it is why published figures for the same named channel vary by more than an order of magnitude: without knowing where the lines sat, there is nothing to compare a benchmark against.
Numr measures the full funnel per client, per project and per time period. The stages, in order:
Stage | What it is |
|---|---|
Records received | What the client's system handed over |
Invalid respondents | Records that cannot be contacted: missing or malformed contact details |
Toxicity rule failures | Records held back by the client's own frequency rules |
Unsubscribe check failures | People who have opted out and must not be contacted |
Send-out failures | Messages the sending system could not dispatch |
Actual send-outs | Records received, less all four exclusions above |
Delivery attempted | Send-outs the channel tried to deliver |
Delivery confirmed | Messages the channel confirms arrived |
Partial | Surveys started and not finished |
Completed | Surveys finished |
Any stage in this table can serve as a numerator, and any stage above it as the denominator. The pair is a decision, and it should be taken and written down before a number exists, not reconstructed afterwards when two reports disagree. That upfront decision is where the variability between response rates comes from, and it is the whole subject of this page. Where Numr holds a counted figure, the page gives it; where it does not, the sentence says so instead of inventing one.
The funnel has two halves that behave differently. Records received down to actual send-outs happens before any customer sees anything: losses there are decided by data quality, business rules and sending infrastructure. Delivery attempted down to completed is where the customer exists: losses there are decided by whether the message arrived and whether the person cared. The halves live in writing that never meets, deliverability material on one side, survey-design advice on the other, and the quoted response rate belongs to both.
Exclusions happen before the denominator
Four kinds of record are removed before anything is sent, so the base being measured is already not the base the client handed over. Invalid respondents cannot be contacted at all: the number is missing or the address malformed. Unsubscribe check failures are people who have opted out and must not be contacted. Send-out failures are messages the sending system could not dispatch. And toxicity rule failures are records held back by the client's own frequency rules.
That last category deserves a definition. A toxicity rule is a client-configured business rule capping how often any individual is invited, to prevent over-surveying; rules are set per project and can be arbitrarily specific. They can be cut by customer type: premium customers no more than once in eight weeks, everyone else once in four. They can be cut by journey: a customer surveyed after a branch visit is not then surveyed about an ATM visit. A well-run programme excludes records this way on purpose, and a programme comparing its rate against a benchmark computed from records received is comparing against a number built on a different base before a single invitation moves.
Two denominators, two rates
Numr's platform computes response rate as answered, meaning partial plus completed, over delivery-confirmed. Where a channel does not capture delivery confirmation, that denominator does not exist and the rate must be computed over actual send-outs instead. The same programme therefore yields two response rates depending on the denominator. The send-out version is never higher, and is lower whenever any send-out fails to confirm: unconfirmed send-outs sit in the denominator contributing nothing to the numerator.
Neither is wrong; they answer different questions. Answered over delivery-confirmed asks: of the people who could have seen this, how many responded? That measures the survey and the moment. Answered over actual send-outs asks: of everything dispatched, how much came back? That blends in the channel's ability to deliver, and the blend is the problem: a blended number cannot be diagnosed. Quote either rate, but say which one it is.
Survey methodology solved this first
The discipline of naming your denominator is older than CX. The American Association for Public Opinion Research publishes Standard Definitions: Final Dispositions of Case Codes and Outcome Rates for Surveys, now in its 10th edition (2023); section 7 defines six response rate formulas, RR1 to RR6. They differ on two questions only. Do partials count as responses? And what happens to cases of unknown eligibility, meaning an address or number where fieldwork ended without establishing whether an eligible respondent was behind it: not a refusal, not a failure, just never resolved.
- RR1: completes only; every unknown case stays in the denominator. The lowest rate.
- RR2: completes plus partials; the same denominator.
- RR3 and RR4: the same numerators; only an estimated share of unknown cases stays in the denominator.
- RR5 and RR6: the same numerators; unknown cases removed entirely. The highest rates.
AAPOR's own framing is that the six run from lowest to highest, separated by nothing except those two choices. It does not ask practitioners to agree on one, only to report which they used, so rates from different surveys can be compared at all.
Numr's answered figure counts partials, the RR2 side of AAPOR's first question. One caveat: AAPOR's six were built for telephone and mail fieldwork, where the dispositions are refusal, non-contact and unknown eligibility. Delivery confirmation is not an AAPOR category; the delivered-versus-sent choice above is Numr's extension of the same discipline to digital channels, not something AAPOR specifies. What carries over is the habit, not the formulas.
A worked example
Every figure in this table is invented for illustration. None is drawn from, scaled from or rounded from any Numr client programme.
Stage | Count (invented) |
|---|---|
Records received | 100,000 |
Invalid respondents | 4,000 |
Toxicity rule failures | 11,000 |
Unsubscribe check failures | 2,500 |
Send-out failures | 500 |
Actual send-outs | 82,000 |
Delivery attempted | 82,000 |
Delivery confirmed | 74,000 |
Partial | 1,900 |
Completed | 7,400 |
Answered is partial plus completed: 9,300. Compute the rate three defensible ways. Over delivery-confirmed: 9,300 / 74,000, or 12.6%. Over actual send-outs: 9,300 / 82,000, or 11.3%. Over records received, which is how a benchmark innocent of exclusions implicitly counts: 9,300 / 100,000, or 9.3%. Three honest rates for one invented programme, the largest over a third higher than the smallest. And 18,000 records, nearly a fifth of what the client handed over, were excluded before anything was sent, most deliberately. A programme that cannot see its own funnel cannot say which of the three numbers it has been reporting to its board.
Reading a low rate diagnostically
A low survey response rate is a symptom with two different diseases, and the single percentage cannot tell them apart. The funnel can.
If the loss sits in the top half, send-outs that never became confirmed deliveries, the problem is the list and the plumbing: stale contact details, a sending identity the receiving side distrusts, authentication failing quietly. The customer never declined anything; the customer never saw anything. Redesigning the questionnaire, sharpening the copy or adding an incentive changes nothing, because the survey is not the thing failing.
If the loss sits in the bottom half, confirmed deliveries that never became answers, the problem is relevance. The message arrived; the person did not act. This is where moment, involvement, speed, language and format live, and where questionnaire and programme work pays.
There is a third reading, upstream of both: if records received collapse into a much smaller pool of send-outs, look at which exclusion is growing. Rising invalids are a data-quality signal from the client's own systems; rising unsubscribes are customers voting on the programme; rising toxicity exclusions are the frequency rules doing their job, or set tighter than anyone remembers deciding. All three are visible before a single customer is involved.
Opposite problems take opposite fixes. A programme that tracks only the blended percentage will spend questionnaire effort on plumbing problems and plumbing effort on questionnaire problems, alternately, until something coincides.
Why people answer surveys at all
People answer when they have something to say and the survey arrives while they still want to say it. Numr's counted experience reduces this to two factors. One is involvement: how emotionally engaged the customer is with the category. A premium vehicle purchase is high-involvement and identity-linked; a flight is a commodity the customer wants to be finished with. The other is the moment: whether something has just happened that the customer has an opinion about. Neither predicts alone, and a weak reading on either caps the rate: high involvement with no moment still gets very little. Industry is a rough proxy for involvement and silent on the moment, which is why benchmark tables are half right and still useless.
The cleanest evidence Numr holds is a paired observation. The same automotive brand, the same customers, the same channel, in the same quarter: surveyed shortly after a vehicle purchase, the programme runs far higher than on a generic relationship ping with no triggering event. Everything except the moment is held constant, and the moment alone moves the rate further than any channel or format change Numr has counted. The figures behind the comparison, and the counted ladder of what programmes achieve, are in What is a good survey response rate, linked rather than repeated here.
The same logic yields a rule Numr states as experience, not as an audit: expect a relationship survey to return a lower response rate than a transactional one. Numr has never observed the reverse. An anniversary survey from a car company gets a lower rate than the same category's post-service survey, because nothing has just happened. The rule follows the event, not the metric, so it applies to effort surveys exactly as to NPS; no figure exists for effort surveys, and none is needed, since nothing about the question changes whether something just happened.
Placing your own programme
Ask two questions about your survey, not your industry. Does the customer care about this category enough to hold a standing opinion? And has something just happened that they have an opinion about? Both yes is the strongest position: the post-purchase automotive survey lives there. Involvement without a moment caps the result: the customer cares, but the anniversary ping gives them nothing to say. A moment without involvement still works, because the event supplies the opinion the category never could: a post-flight survey will outdo a survey about the airline in general. Neither, and no channel or format choice will rescue the ask. Decide which position a survey occupies before pulling levers.
The four levers, in order of impact
Speed, then language, then channel, then format. That ordering is Numr's house standard, and all four sit underneath involvement and moment.
Speed
The sharpest lever, and the one most platforms fail. Surveys sent within a couple of hours of the event do best by a wide margin. Numr holds no signed figure for hours-since-event and will not invent one. The mechanism does not need a decimal. The moment decays. A post-service survey answered while the visit is still the most recent thing in the customer's day is a different proposition from the same survey three days later, when the opinion has cooled. For the same reason, this page offers no best day of the week to send: the published claims contradict one another outright, and Numr has no signed data of its own.
Language
Offering more than one language helps. No figure exists, and the direction is enough to act on. A customer invited in the language he actually thinks in is being asked a smaller favour, and in multilingual markets a single-language invitation writes off part of the base before the moment or the channel gets a say.
Channel
Channel materially changes response rate; the ordering is WhatsApp, then SMS, then email. For South Asia that is a counted, measured Numr figure. Beyond South Asia it is Numr's read of the pattern: it holds across the rest of Asia and in South America; in North America WhatsApp is not that popular; in Europe it is gaining ground but not there yet. The test is penetration, not availability: the ordering applies where WhatsApp is the everyday channel people already live in, because response follows whether a channel is part of the person's day. That is why a channel table computed in one market is worthless in another, and why Numr publishes no channel table and no absolute figures. Channel remains a marginal lever next to the moment.
Format
A chat-style survey pulls two to three percentage points more response than a form. The figure is counted and it is exactly that size: percentage points, not a multiplier. On a large base that is real volume, and it is still the smallest of the four levers, which is the category in miniature: the most demonstrable lever is the least important one.
Reminders
Whether reminders pay is answerable from records already held, not from a benchmark. Numr records every invitation sent, which invitation each response came from, and when each reminder went out. Completion by reminder wave is therefore a query against records already held, reportable on a client's own sends: your figure, on your base, in your period.
There is no standing Numr figure for how much each wave adds, and this page will not imply a typical shape. The curve is knowable per programme and unknown in general; a programme weighing a second or third reminder should run the query and decide on its own curve.
What a low rate costs, and what a forced high one hides
A low response rate hurts CX decisions in two ways, and they run on different arithmetic: precision is governed by how many answers you have, representativeness by who gave them. The first cost is width, and it follows the count, not the rate. Fewer answers mean wider uncertainty on every cut, and the cuts that drive action, this branch, that model, last week, go noisy first. A low rate on a very large base can still leave the counts comfortable; the score can look stable while everything beneath it bounces.
The second cost is representativeness, and it belongs to the rate: volume does not cure non-response bias. The people who answer skew to the involved and the recently triggered. If willingness to respond correlates with what the survey measures, the score drifts from the base's true position, and more responses from the same slice add more of the same skew. NPS is a score on a -100 to +100 scale, and a score computed on the engaged fraction of a base is a real number about the wrong population. Weighting repairs skews you can observe, age, region, product line; it cannot repair the one driving response, because the silent customers' engagement was never measured. The lower the rate, the more room the bias has: the rate bounds how wrong the score can silently be.
Chasing the rate can also fail in the opposite direction. A rate inflated by pressure, repeated reminders, prize draws, pleading copy, changes who answers and how: the recruits are the previously unwilling, and answers given to end a conversation are worth less than answers given to have one. The object was never a bigger numerator. It is answers from a slice that resembles the base, given by people who meant them. The funnel keeps this honest: response extracted rather than earned has to surface somewhere, and the stages pressure touches, the partial count and the opt-out exclusions, are where to look. A partial counts in the rate by the stated convention and still works as a warning: a started survey is a response, and a growing share of starts that never finish is the questionnaire talking.
How weak responses are caught
This is detection, not improvement: what follows flags suspect responses, it does not earn better ones. Numr runs two checks.
The first is speed. The platform captures the total duration of every survey and computes from it the average response time per question, and the flag fires on that per-question average, not on the total. The distinction matters because total duration on its own tells you nothing: a long questionnaire answered carefully and a short one answered inattentively can produce the same number. The per-question average can be compared across questionnaires of different lengths, which is what makes it usable as a flag at all. If the average drops too low, the survey is flagged as answered too quickly. No number is published for too low.
The second is straightlining, and Numr's version is cut to the two shapes CX questionnaires actually take. In a rating-only questionnaire, where every sub-attribute question is a rating, the flag is every question receiving the same rating. In a drill-down questionnaire, where each option opens a further set of choices, the flag is only the first choice being selected at every level of every drill-down. That second pattern is worth spelling out, because a generic straightlining check cannot see it. The answers are not identical; they are all first.
A flagged response is flagged. It is not removed by default, and the difference is not administrative. Silently dropping flagged records changes the denominator without saying so, which is precisely the defect this document exists to attack. A flag is a signal, not a verdict. The decision to exclude a response is a decision about the data, it belongs to whoever owns the programme, and it should be made deliberately and written down, not performed quietly by a vendor's cleaning routine.
Two judgements matter more than any threshold. Set your thresholds before you look at the results, because a threshold chosen after you have seen the scores is a way of choosing the scores. And if a large share of responses is failing the checks, suspect the survey before the respondents.
Frequently asked questions
What is a good survey response rate?
No single number exists; any table offering one answers half the question. Involvement and the moment drive the rate, and a weak reading on either caps it, whatever the channel. The counted figures for what programmes achieve are in What is a good survey response rate.
What is the difference between response rate and completion rate?
Response rate is replies over invitations. Completion rate is finished surveys over surveys started. They fail for different reasons: a programme can hold a flat response rate while completion collapses, which usually means the invitation still works and the questionnaire does not. Track both, separately.
Do reminders increase survey response rates?
Usually, but no general figure exists for how much. On a Numr programme, completion by reminder wave is a query against records already held, so the question is answered with your own figure on your own sends.
What is the best day or time to send a survey?
Within a couple of hours of the event, whichever day that is. Numr holds no signed figure for hours-since-event and no day-of-week data; the published day-of-week claims contradict one another. Speed to the moment is the sharpest lever; a survey arriving while the event is fresh makes the day-of-week question irrelevant.
Which channel gets the best survey response rate?
WhatsApp, then SMS, then email, wherever WhatsApp is the messaging channel people already live in. Counted for South Asia; observed, not counted, across the rest of Asia and South America; in North America WhatsApp is not that popular, and in Europe it is gaining ground but not there yet. Still no channel table and no absolute figure by channel, and channel stays marginal next to the moment.
Can a low response rate make my NPS wrong?
Yes. Respondents skew toward the involved and the recently triggered, so the computed score can sit some distance from the full base, and adding volume from the same slice does not close the gap. NPS is a score on a -100 to +100 scale; the lower the response rate, the wider the room between measured and true.
Why does my relationship survey get a lower response rate than my transactional one?
Because nothing has just happened. A transactional survey arrives on the heels of an event the customer has an opinion about; a relationship survey asks about nothing in particular. Numr has never observed the reverse. The rule follows the event, not the metric, so it holds for effort surveys as for NPS.
The funnel shows thousands of records excluded before anything was sent. Is that a problem?
Not by itself; several of those exclusions are the programme working as designed. Toxicity exclusions are deliberate protection against over-surveying; invalid records are a data-quality signal for whoever owns the source system; a growing unsubscribe count is the one to watch: customers rendering a verdict on the programme. The problem is a programme that reports a response rate without knowing its exclusions exist.