What Should I Track Once the Survey Is Live?

The four things to monitor once a survey goes live: delivery rate, segment distribution, response rate and drop-off. With thresholds and when to stop reminders.

· Last updated August 21, 2026

TL;DR

  • Monitor four things in order: delivery, segment distribution, response rate, and completion.
  • Delivery is a technical health check. Below 95 percent delivered for email, stop and fix the list or the sender setup before reading any results.
  • Distribution matters more than volume. A representative 12 percent beats a lopsided 30 percent, because a skewed sample produces a confidently wrong score.
  • Two reminders is usually the ceiling. Stop when new responses stop changing the score.

A live survey is an operational system, not a completed task. Most of the damage done to CX data happens in the days after launch, quietly, while everyone waits for results. These are the four checks that catch it.

1. Delivery rate

What it is: invitations successfully delivered, divided by invitations sent. Delivered, not opened.

Threshold: 95 percent or better for email. SMS should sit higher, around 97 percent, because the failure modes are narrower.

Check it: within the first hour of the send, then daily.

What failure looks like: a sudden drop usually means list decay, a bad export, or a sender reputation problem. Look at the bounce split. High hard-bounce counts point at stale contact data, often a segment that has not been cleaned in years. High soft bounces or blocks point at your sending domain: authentication records, a shared IP with a poor reputation, or a spam filter reacting to the send volume.

Why it comes first: delivery failure is rarely random. It concentrates in older records, particular corporate domains, or a single region, which means a delivery problem quietly becomes a sample problem. Every downstream number inherits the bias.

Always check delivery by domain and by segment, not just in total. A 96 percent overall delivery rate can conceal one corporate domain blocking you entirely, which silently removes an entire customer type from your results.

2. Segment distribution

What it is: the mix of your responding sample against the mix of the population you invited, across the dimensions that matter: region, tenure, product, channel, value tier, language.

Threshold: each key segment's share of responses should sit within a few percentage points of its share of the invited base. Set the tolerance before you launch so you are not negotiating with yourself afterwards.

What failure looks like: one segment supplying most of the responses. Long-tenure customers almost always over-respond. Digital-native segments over-respond to email. High-value accounts under-respond because they are busier. Each skew bends the score in a predictable direction, and none of it is visible in a headline number.

What to do: first try to fix the collection, with a targeted reminder or a second channel for the under-represented group. Only weight the data as a fallback, and disclose the weighting whenever the score is reported. Weighting repairs the arithmetic, not the underlying silence.

A skewed sample does not produce a vague answer. It produces a precise, confident, wrong one. The teams most likely to act on bad data are the ones whose response rate looked healthy.

3. Response rate

What it is: completed responses divided by invitations delivered. Not sent. Publish the denominator alongside the number, because switching denominators is the easiest way to make a response rate look better than it is.

Threshold: it depends heavily on channel, relationship and how recently the interaction occurred, so treat external figures as context rather than targets. A 2022 meta-analysis of 1,071 online survey response rates in published research found an average of 44.1 percent, and also found that sending to more people did not raise the rate. A clearly defined, refined population did (Wu, Zhao and Fils-Aime, Computers in Human Behavior Reports, 2022). Academic populations are not customer populations, so the useful takeaway is the direction, not the number: targeting beats volume.

Check it: at 24 hours, 72 hours, and at close. The 24-hour figure is your best early predictor. Transactional surveys collect most of their responses in the first day, and a weak first day rarely recovers.

What failure looks like: an unusually low first-day rate points at the invitation, not the survey. Subject line, sender name, timing, or an invitation that reads like marketing. A rate that starts well and stalls points at the survey itself, which is the next check.

One important caveat. A higher response rate is not automatically better data. Hendra and Hill examined survey data from a large-scale evaluation and found little relationship between response rates and nonresponse bias, noting that chasing a high rate lengthens the field period in ways that create their own measurement problems (Evaluation Review, 2019, 43(5), 307 to 330). Representativeness is the goal. Response rate is a proxy for it, and an imperfect one.

4. Completion and drop-off

What it is: completed responses divided by surveys started, plus the drop-off curve showing where people abandon.

Threshold: 85 percent or better completion for a short transactional survey. Longer relationship surveys will run lower, but a completion rate under 70 percent means the instrument is asking too much.

Check it: question by question. The aggregate completion rate tells you that people are leaving. The per-question curve tells you where, which is the only version you can act on.

What failure looks like: a single question causing a visible cliff. The usual culprits are a mandatory open-text field, a question asking for information the customer does not have to hand, a grid that collapses on mobile, or a demographic block placed too early. Move it, make it optional, or cut it.

Also watch time to complete against your estimate. If the real median is meaningfully longer than what the invitation promised, the drop-off is a broken promise rather than a bad question.

"The first thing we look at on a live survey is not the score. It is whether the responses look like the customer base. We have seen programmes run for a year on a sample that quietly excluded one corporate email domain, so an entire customer type was missing from every board report. Delivery by domain would have caught it in an afternoon." Samudra Gupta, CTO, Numr

When to stop sending reminders

Reminders lift response rates and then start costing more than they return. Three rules keep that trade honest.

Cap at two reminders. One at roughly 48 to 72 hours, a second near the close of the field period. A third mostly annoys people who have already decided not to answer, and it disproportionately reaches your most loyal customers, who are also the most likely to be on every other list you run.

Stop when the score stops moving. Track the running score as responses arrive. Once a new batch of responses no longer shifts the score beyond ordinary noise, extra volume is buying precision you do not need at a cost in goodwill you do.

Stop when reminders stop fixing the skew. The strongest reason to send a reminder is a missing segment, so target it there. Once that segment is proportionally represented, the reminder has done its job.

Then close the field on schedule. An indefinite field period makes period-over-period comparison meaningless, and comparability is the whole reason you are tracking a metric in the first place.

After the field closes

Record the instrument alongside the result: question wording, scale, trigger, field dates, denominator, reminder schedule, and any weighting. When a score moves next period, the first question is always whether the experience changed or the measurement did. Without that record you cannot answer it, and a movement you cannot explain is a movement nobody will act on. This is also why unexplained score changes should be checked against the instrument before the experience, as covered in what is a good NPS, CSAT or CES benchmark.

Frequently asked questions

What is a good delivery rate for a survey invitation? 95 percent or better for email, and around 97 percent for SMS. Below that, fix the list or the sending setup before reading the results, because delivery failures cluster in particular segments and turn into sample bias.

Should response rate be calculated on invitations sent or delivered? Delivered. Using sent understates the rate and hides delivery problems. Whichever you use, publish the denominator so the number can be compared over time.

Is a higher response rate always better? No. Hendra and Hill found little relationship between response rate and nonresponse bias in a large-scale evaluation. A representative sample at a modest rate is more trustworthy than a skewed sample at a high one.

How many reminders should I send? Two at most, ideally targeted at the segments that are under-represented rather than blasted to everyone who has not replied.

What completion rate should I expect? 85 percent or better for a short transactional survey. Under 70 percent indicates the survey is too long or contains a question people will not answer.

How quickly should I check a live survey? Delivery within the first hour. Response rate at 24 hours, which is your best early predictor. Distribution and drop-off daily until the field closes.

What if one segment refuses to respond at all? Try a different channel or a different invitation for that group before resorting to weighting. If it stays silent, say so explicitly when reporting the score rather than letting the gap pass unnoticed.

For related reading, see How to Improve Response Quality, Not Just Response Rate, How to Design a CX Survey That Customers Actually Complete, and Exploratory vs Descriptive vs Predictive Survey Research.

Share

See Numr CXM answer your next question.