How Many Survey Responses Do You Actually Need?

Abstract Numr artwork: one measured scoop of coffee beans held above a grid of open, full bags.

Survey sample size: the short answer

How many survey responses do you need? For a single yes-or-no question, at 95% confidence and a margin of error of plus or minus 5 points, the survey sample size is 385. That figure is correct, every calculator produces it, and it is the answer to a question you probably are not asking.

Every survey sample size formula assumes three things: that you are estimating one proportion, that the population is large enough to treat as infinite, and that the people who answered are a fair draw from the people you asked. A customer experience survey usually breaks all three. Breaking the infinite population means you need fewer than 385. Breaking the single proportion means 385 is not enough, in three separate ways. Breaking the fair draw is not a sample size problem at all, and it is the one that decides whether the data is worth reading.

That last point is why "statistically valid survey response rate" is a term that cannot be answered with a number. Validity is not a quantity of responses. It is a property of who answered.

Where 385 comes from

The formula is short and worth seeing, because everything after this is a departure from it.

n  =  z² × p(1−p) / e²

`z` is 1.96 for 95% confidence. `e` is the margin of error you will accept. `p` is the proportion you expect, and using 0.5 is the cautious choice because that is where a proportion varies most, so it produces the largest sample any answer could require.

Run it at three common targets:

Margin of error

Responses needed

plus or minus 10 points

97

plus or minus 5 points

385

plus or minus 3 points

1,068

Halving the margin of error from 10 to 5 costs four times the responses, and going from 5 to 3 costs nearly three times again. Precision is bought at an accelerating price, which is why a demand for plus or minus 3 should be questioned rather than met.

Departure one: your list is finite, so you need fewer

The formula assumes you are drawing from an endless supply of people. A pollster sampling a country is close enough to that, and the assumption costs nothing. You are drawing from a customer list, which is a fixed group and often a small one, and there the assumption makes you collect responses you did not need.

Here is why. Suppose your list is 500 people and 400 of them answer. You have not sampled that group. You have very nearly counted it, and what is left to estimate is the 100 who did not answer, not the whole 500. The formula cannot see this. It asks for 385 whether your list holds 500 people or five million.

The fix has a name, the finite population correction, and it does exactly one thing: it tells you where you can stop.

n_adjusted  =  n / (1 + (n − 1) / N)

where `N` is the size of your list. The first row below is the textbook answer, the one every calculator gives. Every row under it is what you actually need.

Your list

at plus or minus 5 points

at plus or minus 3 points

no limit (the textbook)

385

1,068

50,000

382

1,045

10,000

370

965

5,000

357

880

2,000

323

697

1,000

278

517

500

218

341

What you are saving is responses, and behind each response is an invitation, a reminder and somebody's time. On a list of 500 you stop at 218 instead of 385. That is 167 fewer people who have to answer, at the same 95% confidence and the same plus or minus 5 points. Nothing is traded away. The 385 was simply the wrong number for a list that size.

How large the discount is depends on what share of your list you are asking, not on how big the list is. Ask most of the list and the discount is large. Ask a sliver and there is barely one.

As a rough guide, the discount is about the share you end up asking.

list of 10,000    ask 370   =  3.7% of it    save 15   =  3.9%
list of 500       ask 218   = 43.6% of it    save 167  = 43.4%

Measure the share against the number you actually ask, not against the textbook 385. On that list of 500 the 385 is 77% of everyone you have, which would have you expecting a discount nearly twice the real one.

That is also why the right hand column saves more than the left. Tightening the target from plus or minus 5 to plus or minus 3 needs more responses, and more responses is a larger share of the same list. On a list of 5,000 you save 28 responses at plus or minus 5 and 188 at plus or minus 3.

So there is no list size above which this can be ignored. There is only a point per target where it stops being worth the arithmetic, and that point moves whenever the target does.

Departure two: every cut spends the 385

Whatever number you set out to collect is where you start. Call it 385. It does not follow you through the report.

The margin of error belongs to a single proportion, and a questionnaire is a stack of them. Every step between collection and the number on the screen takes something off the base, and the margin of error widens as it goes. Nothing about this is hidden. It is simply never revisited: the 385 was agreed at the top of the project, and nobody carries the margin of error down the chain behind it.

Follow one number down the chain. You collect 385. Skip logic and non-applicable options mean 240 of them are shown question 14 and answer it. You cut that by region and read 80. Then you compare that region to another one.

385 collected, all answered, split 50/50      +/- 5.0 points
240 answered question 14, split 50/50         +/- 6.3 points
80 in one region, split 50/50                 +/- 11.0 points
the gap between two regions of 80             +/- 15.5 points

Three cuts, and the margin of error has tripled. No single cut looks unreasonable while you are making it.

The last row is the one that catches people, because comparing is a different calculation from reporting. The margin of error on a difference carries the uncertainty of both estimates, not of either one. Two regions eighty apiece can differ by fifteen points and be indistinguishable.

The error runs the other way too. Two estimates whose intervals overlap are not therefore the same. Those same two regions, twenty points apart, have intervals that still overlap by about two points, and the gap is significant at z = 2.53. Non-overlap is the stricter test: it implies z above 1.96 in every case, and 2.77 or more when the two groups are the same size and split much the same way. So eyeballing the overlap never invents a gap. It quietly misses real ones, like this pair.

Which is why "is the base big enough" is the wrong question to ask of a report. The right one is what the margin of error is at the point you are reading, after everything upstream has taken its cut. And if comparing segments is the point of the exercise rather than a thing you do afterwards, size for it: getting the margin of error on a difference down to ten points takes about 193 in each group, which is a different survey from one built to put 385 in the headline.

Departure three: NPS is a difference, and differences are noisier

This one matters most, because it applies to the metric most of these surveys exist to produce.

NPS is not a proportion. It is the promoter percentage minus the detractor percentage, and subtracting two uncertain numbers gives a result more uncertain than either. At the same sample size, a net score is wider than either proportion it is built from, on any base containing both promoters and detractors. The two are equal only in the degenerate case where one of the groups is empty.

Standard error of a net score  =  √( (p + d − (p − d)²) / n )

At 385 responses, where the margin of error on a proportion is plus or minus 5 points:

Promoters / Passives / Detractors

NPS margin of error

60 / 30 / 10

plus or minus 6.7 points

50 / 30 / 20

plus or minus 7.8 points

40 / 30 / 30

plus or minus 8.3 points

50 / 0 / 50

plus or minus 10.0 points

The last row is the worst case, a fully polarised base with nobody in the middle, and it is exactly twice the margin of error on a proportion at the same n. The more polarised your customers, the less precisely you can measure the gap between them, which is an unwelcome property in a metric designed to detect polarisation.

Now apply departure two to it, because the thing you actually want to read is a change. A quarter-on-quarter comparison is a difference between two net scores, so both quarters contribute uncertainty. On a base like 50 / 30 / 20 with 385 responses in each quarter, the margin of error on one quarter's NPS is 7.8 points; on the change between the two, 11.0.

A five point move, on that base, cannot be distinguished from no change at all. Neither can a ten point one. That is not a reason to stop reading the trend, and departure four is about what to do instead, but it is a reason to stop writing the quarterly commentary as though a five point move were an event.

Scales and multi-select: two question types with their own arithmetic

CX questionnaires lean on five point scales, seven point scales and select-all-that-apply lists. The formula still touches both, but not evenly. Report a scale as a mean and its arithmetic is replaced entirely. Report a multi-select and each option is an ordinary proportion, until you compare two of them, at which point the formula needs a number the two percentages do not contain.

If you report the mean of a scale

Report a top-2-box percentage and you are back to a proportion, so everything above applies unaltered. Report the mean and `p(1−p)` drops out of the arithmetic entirely. What replaces it is the spread of the answers themselves.

margin of error on a mean  =  z × s / √n

where `s` is the standard deviation of the answers you collected. This is the one place on this page where you cannot pick a safe default the way p = 0.5 is safe for a proportion. You have to supply `s` from your own data. You can bound it, though: on a five point scale the widest possible spread is half your respondents at 1 and half at 5, which gives s = 2.0.

At 385 responses on a five point scale:

What you measure

Margin of error

the mean, illustrative spread (s = 1.1)

0.11 scale points

the mean, worst case spread (s = 2.0)

0.20 scale points

the same data as top-2-box at 60%

4.9 percentage points

The first and third rows are the same 385 people. The mean is measured more finely because it uses every answer, while collapsing to top-2-box discards the difference between a 1 and a 3. Top-2-box is easier to explain and easier to act on. The price is precision.

Before the 0.11 reassures you, apply departure three. A change between two quarters carries a margin of error 1.4 times wider, so 0.16 scale points. A quarter-on-quarter move of a tenth of a point cannot be distinguished from no change. Getting the margin of error on a change down to 0.10 takes 930 responses in each quarter.

If you ask select all that apply

A multi-select gives you one proportion per option, each computed on the full base. The percentages sum to more than 100% and that is fine. They are not shares of a whole.

Two things go wrong, and only one of them is about sample size.

The tail is much less precise than it looks. At 385 responses:

Option chosen by

Margin of error

95% range

62%

4.8 points

57% and 67%

41%

4.9 points

36% and 46%

22%

4.1 points

18% and 26%

8%

2.7 points

5% and 11%

The 8% option carries the smallest margin of error on the list in points, and by far the largest relative to itself. Its true value is somewhere between 5% and 11%, a factor of two. Ranking the bottom half of a multi-select list is not something 385 responses will support.

Comparing two options needs the overlap, and the overlap is not in the two percentages. Say one option is chosen by 41% and another by 35%. The margin of error on that six point gap depends entirely on how many people chose both:

nobody chose both                                +/- 8.7 points
the two choices independent of each other        +/- 6.8 points
everyone who chose the 35% also chose the 41%    +/- 2.4 points

Same two percentages, same 385 people, and the margin of error swings by a factor of nearly four. Without the cross-tab you cannot say whether a six point gap is real, and no sample size answers it, because it is not a sample size question.

What you need is one count: how many respondents chose both options. Any survey tool can produce it, and it is the difference between knowing which of your gaps are real and guessing at all of them.

The top row is worth recognising. Nobody choosing both means the options are mutually exclusive, and the margin of error there is the net score formula from departure three, unchanged. Promoters and detractors are mutually exclusive options, which is why NPS behaves the way it does.

Departure four: no sample size repairs a self-selected sample

Everything so far has been arithmetic. This one is not, and it is the one that decides whether the exercise was worth doing.

Every formula on this page assumes the people who responded are a fair draw from the people you invited. In a customer survey they are not. People who answer digital surveys are disproportionately those with something to say, which hollows out the middle and leaves you with the pleased and the annoyed. Collect the same programme by telephone and courtesy pushes the score up instead. Neither set of respondents is the customer base.

Increasing the sample does not fix this. It measures the same skewed population more precisely. Ten thousand responses from people who chose to answer will give you a beautifully tight estimate of what people who choose to answer think.

This is where departure three resolves rather than contradicts. Departure three says a small move is inside the margin of error, so it cannot be read as a result. Departure four says the level was never accurate anyway, because the sample is skewed. Both point at the same discipline. Real is not the objective. Movement is. Hold the method still, ask whether this quarter is above the last one measured the same way, and a consistent bias stops mattering because it is present in both readings. Then give the movement a base large enough or a run of quarters long enough to clear the margin of error on a difference, rather than declaring a result from a single step that the arithmetic cannot support. A measure that is slightly wrong in a stable direction is more useful than one that is truer and changes between quarters.

So what do you actually do

Start from the decision the number has to support. If a five point move would not change anything you do, you do not need a margin of error tight enough to detect one, and you should not pay for 1,068 responses to find out.

Then four questions, in this order.

  1. Is the headline a proportion or a net score? The net score is wider. At 385 responses the margin of error on a proportion is plus or minus 5.0 points; on an NPS with a 50 / 30 / 20 base it is plus or minus 7.8.
  2. Are you reporting a level or a change? The margin of error on a change is wider again, by a factor of 1.4. At 385 the proportion's 5.0 becomes 7.1, and the NPS 7.8 becomes 11.0.
  3. How big is the list? Below about 2,000, at the plus or minus 5 target, the discount for a finite list is worth having: 62 fewer responses at a list of 2,000, 107 fewer at 1,000, 167 fewer at 500.
  4. Will you compare segments? If yes, size each segment for the comparison rather than dividing the headline among them. About 193 per group brings the margin of error on a gap down to ten points, so the total scales with the number of segments.

Three cases, worked.

A B2B programme. 900 customers, one headline satisfaction number.

target                        +/- 5 points
uncorrected requirement       385
corrected for a list of 900   270
at a 30% response rate        invite 900

At 30% the sums close exactly: 270 expected, 270 required. A few points lower and even the full census misses the target, since 25% of 900 is 225 responses and a margin of error of plus or minus 5.7. A few points higher and you could spare a hundred invitations, which is not a plan. At this size there is no sampling plan, only a census with a response problem.

A consumer programme reporting four regions.

headline only              385
four regions, compared     193 each
total                      772

Comparing costs twice what reporting costs. Decide before fielding, because you cannot go back and top up a region's base afterwards.

A quarterly NPS tracker on a 50 / 30 / 20 base.

margin of error on the change, at 385 per quarter    11 points
responses needed to read a 5 point move             1,875 per quarter

Trackers routinely run at a few hundred responses and discuss five point moves every quarter. Those two facts are difficult to hold at the same time.

Then work backwards to invitations, using a response rate observed at a comparable moment rather than a generic one. That ratio decides how many people you have to contact.

Two habits

Publish the margin of error, not just the base. A base leaves the reader a conversion that nobody performs. The conversion is one line: a 40% score on a base of 80 is 1.96 × √(0.4 × 0.6/80), or plus or minus 10.7 points. A margin of error is readable on sight, and it is what the decision turns on.

Do not confuse suppression with significance. Low base suppression on a dashboard is a governance control, not a significance test: it decides what a viewer should be shown. Significance decides what the data supports. Numr draws the line explicitly. A driver reaches the priority matrix when its regression coefficient clears significance, and there is no minimum base for visibility, because the p value governs it.

---

Frequently asked questions

How many responses do I need for a statistically valid survey?

For one question at 95% confidence and a margin of error of plus or minus 5 points, 385 responses. You need fewer if the list is finite: 218 on a list of 500, 278 on 1,000, 357 on 5,000. You need more if your headline metric is a net score like NPS, which can carry twice the margin of error at the same sample size. The word "valid" is doing unearned work in that question, though. A sample size buys precision and nothing else. A survey answered by 385 self-selected people is precise and unrepresentative at the same time.

What is the formula for sample size?

n = z² × p(1−p) / e², where z is 1.96 for 95% confidence, e is the acceptable margin of error as a decimal, and p is the expected proportion. Using p = 0.5 gives the largest sample any answer could need, which is why calculators default to it. If you have a prior estimate, use it: at p = 0.9 the same margin of error needs 139 responses instead of 385.

Does NPS need a bigger sample than other metrics?

Yes. NPS subtracts one proportion from another, and the uncertainty of a difference is larger than the uncertainty of either part. At 385 responses the margin of error on a proportion is plus or minus 5 points, while on NPS it runs between roughly 6.7 and 10 points depending on how polarised the base is. The worst case, an even split of promoters and detractors with no passives, is exactly twice as wide.

Is 100 responses enough?

For one question where you can live with plus or minus 10 points, 97 responses is the arithmetic answer, and on a list of 300 the finite population correction brings the requirement down to 73. On a list of 100 or fewer the arithmetic is moot: survey everyone, and the only error left is who declined. What 100 responses will not do is support a comparison. At that base an NPS reading carries roughly plus or minus 15 points on a typical 50 / 30 / 20 split, and 19.6 on a fully polarised one, so nothing short of a very large quarter-on-quarter move registers at all, and any subgroup cut out of the 100 is unreportable.

Does a bigger sample fix a low response rate?

No. A low response rate is a warning about who answered, not about how many. Increasing the sample measures the same self-selected group more precisely. The defence is to hold your method constant and read the movement rather than the level.

Share

See Numr CXM answer your next question.