Transactional and Relationship NPS: Why the Gap Exists, and How to Link Them
TL;DR
- Transactional and relationship NPS are not two measurements of one thing. One captures a customer in the moment, the other captures their reflection over time.
- There is no standard gap between them and there cannot be one. The gap is a product of which touchpoints you measure, how fast each feeds through, and brand effects outside the programme entirely.
- The gap has no fixed direction. Transactional sitting above relationship is as normal as below, and the direction is not a diagnosis.
- Closing the gap should not be a goal. Two surveys that converged would have stopped measuring two different things.
- The two can be connected, and Numr calls this linkage analysis. It needs a rating question about each touchpoint in both surveys, though not the same question and not the same scale. It then measures how long each touchpoint takes to feed through, works out how much each one counts, and uses the earlier transactional reading to forecast the recommendation rating before the next relationship wave is fielded.
The question is almost always the wrong one
A client asking about the gap between their transactional and relationship NPS is usually asking which number is wrong.
Neither is. The gap is what you should expect from two instruments pointed at two different things, and the useful conversation starts once that is out of the way.
A transactional survey asks a customer about something that has just happened to them. A relationship survey asks what they think of you, in general, now. Those are different questions producing different emotions, and the numbers are not two estimates of a single truth.
There is now direct evidence for this. In a field study of more than 5,000 hotel stays collected over five years, published in the Journal of Business Research in 2024, Heath McCullough and colleagues introduced a construct they call retrospective delay: the time between an experience happening and the customer evaluating it. They found that the delay changes the answer, and that first impressions become more influential as the delay grows. Their own conclusion is that measurement timing "could explain some conflicting findings in the services literature."
Which is the point, in plain terms: two surveys asked at different distances from the event should not be expected to agree. That is a property of measurement, not a defect in your programme.
A note on the labels, because they mislead
"Transactional NPS" and "relationship NPS" name two kinds of survey, not two kinds of NPS. tNPS is a transactional survey, sent after something happened. rNPS is a relationship survey, sent on a cadence. The metric is the same metric; what differs is when it is asked and about what.
This matters more than a naming quibble, because the labels invite people to think of them as two scores that ought to reconcile, and the rest of this document is about why they do not. It also matters for the model described later, which does not operate on either headline score. It operates on rated questions inside the two surveys.
What moves the relationship score
Your relationship score comes from your relationship survey. It is measured, not assembled, and nothing below changes that.
What moves it is your touchpoints. Each one contributes, each one contributes a different amount, and each one arrives on its own schedule. That is also what makes a predicted relationship score possible, which is what linkage analysis produces later in this document: a forecast of the number the relationship survey will report, built from touchpoint performance that has already happened.
Two things follow, and almost every misreading of the gap comes from missing one of them.
Each transactional score feeds in with its own lag. The relationship score therefore responds to transactional movement later than the transactional score does, and different touchpoints feed through at different speeds. Onboarding and a complaint resolution do not reach the relationship view on the same timetable.
Each touchpoint carries its own weight. A touchpoint that almost every customer passes through, badly, will move the relationship score more than one that a tenth of customers touch. Those weights are not equal, and they are not guessable. They have to be measured on your data.
That is the mechanism behind the gap. It is also why a relationship programme can read flat while the journeys underneath it are already improving. A client who has fixed something real and sees no movement in the headline number is not being told the fix failed. They are being told it has not arrived yet.
There is no standard gap, and there cannot be one
The gap is a product of which touchpoints a client measures, how those scores feed through with their own lags, and brand effects sitting outside the programme entirely.
Two clients with identical journeys will not have the same gap. Anyone quoting a standard delta is describing one programme and calling it a rule.
What is measurable is the gap in one programme over time, on its own data. That is the same position that applies everywhere else in measurement: hold the method constant and watch your own number.
The gap has no fixed direction
This is the half almost nobody expects.
Sometimes the transactional score sits below the relationship score. Sometimes it sits above. Both are normal.
A customer can find individual interactions effortful while still liking the product, the brand or the price, so the overall view survives the friction. Equally, every interaction can be handled well while the customer carries something none of those interactions caused.
Those are illustrations of how either direction can arise. They are not readings of what a direction means, and the distinction matters, because at least three mechanisms produce the same observable gap:
- Feed-through lag. The relationship score moves later, so a gap may simply be a timing artefact of a programme that is already improving.
- Coverage. A client measuring three journeys out of ten has a transactional average describing a third of the customer's experience. The other seven journeys are in the relationship score whether they are measured or not.
- Everything outside the journeys. Pricing, brand, a competitor's launch, a news cycle. No CX programme reaches these, and they sit in the relationship number.
Which of the three is operating is a question for that client's own data. Anyone who tells a client which way round the gap ought to be is describing one programme rather than stating a rule.
Closing the gap should not be a goal
A client who sets gap reduction as a target has set the wrong target.
A transactional survey should never be identical to a relationship survey. The in-the-moment impression does feed the over-time impression, but it feeds in with a delay, and by the time it arrives the emotion may be weaker or stronger than it was.
A programme where the two converged completely would have stopped measuring two different things. The goal is not to make them agree. It is to know what each one is telling you, and to watch each of them move against itself.
If a client wants to act, the thing to act on is the relationship score itself, never the distance between the two numbers. What moves it is improving the transactional experiences that feed it, and then allowing the feed-through.
Linkage analysis: connecting the two properly
The scores cannot be compared directly. A transactional programme reading, say, in the low fifties and a relationship programme in the high seventies are not two readings of one quantity.
The level of the gap is uninterpretable, but changes in it are not. Across programmes the difference between two such numbers means nothing, because the instruments, the coverage and the populations differ. Within one programme, held to a constant method, a gap that moves is telling you something, and the three mechanisms above are the candidate explanations for the movement. That is the distinction between a number worth watching and a number worth comparing.
They can, however, be linked. Numr calls this linkage analysis, and it has been run across several client programmes in different sectors.
The bridge: a rating question about the touchpoint, in both surveys
The design choice that makes everything else possible is this. For every touchpoint that runs its own transactional survey, the relationship survey also carries a rating question about that touchpoint.
They do not have to be the same question, and they do not have to be on the same scale. The transactional survey might ask how easy the interaction was, on whatever scale that programme uses. The relationship survey might ask how satisfied the customer is with that touchpoint overall, on 0 to 10. Both are rating questions about the same touchpoint, and that is the requirement.
Two reasons the scales can differ, which are worth understanding rather than taking on trust. The lag is read from the shape of two trends over time, not from their levels, and shape survives a change of scale: a series that rises and falls in a particular pattern still rises and falls in that pattern whether it runs 1 to 7 or 0 to 10. And the regression that produces the weights runs entirely inside the relationship survey, so only that survey's scale ever enters a coefficient.
What does not work is a yes or no on either side, or a categorical answer. There is nothing to trend and nothing to regress. Rating questions are the requirement; matching them is not.
This is the step most easily skipped, and skipping it is what makes the two scores impossible to connect later.
Step one: measure the lag, do not assume it
Trend the touchpoint rating as it is answered in the relationship survey. Trend the touchpoint rating from the transactional survey. Compare the shapes of the two series over time.
In practice you slide one line along the other until the two shapes line up. How far you had to slide it is the lag for that touchpoint, measured rather than assumed. It is not a rule of thumb, and it is not the same for every touchpoint or every client. A fast-resolving process and a long institutional one feed through on different timetables, and the difference between them is not small.
The obvious practical question is whether you have enough history to see it. What matters is the number of data points, not the number of months, and a transactional programme collecting continuously can be cut into weekly points, which gives you fifty-two in a year rather than four.
But you are sliding two series, so the sparser one sets the resolution, and the sparser one is almost always the relationship survey. Cutting the transactional stream finer does not help if the thing you are sliding it against reports four times a year. Four points will not identify a lag of any length worth knowing about, and this is the constraint that actually bites. The relationship survey has to run at a comparable cadence to the transactional programme.
Which sounds like a demand for more surveying and is not. The move is to field the relationship survey continuously instead of in waves, by triggering it on something each customer has that is spread across the calendar. Two shapes that work:
- An anniversary trigger. Each customer is surveyed at a fixed point relative to their own purchase or joining date. Because customers joined on different days, the programme fields every week while any individual customer is asked once a year. One automotive programme runs exactly this, and the anniversary survey doubles as the relationship survey rather than sitting alongside one.
- A renewal trigger. Each customer is surveyed at a fixed interval before their renewal date, with the data arriving as a monthly file. One B2B infrastructure programme surveys three months ahead of renewal on that basis.
Both give a rolling relationship series dense enough to slide against a transactional one, without asking any customer to answer more often than they would have anyway.
Numr's normal recommendation is at least a year of data before attempting a linkage, and more where either side is sparse. Which is the argument for designing this in before anyone asks for the analysis: a rating question per touchpoint in the relationship survey, and a relationship survey that fields continuously. A programme that has neither cannot link at all, however good its data otherwise.
Step two: weight each touchpoint by regression
Regress the touchpoint ratings against the recommendation question, both drawn from inside the relationship survey.
That produces a weight for each touchpoint, called a beta: how strongly that touchpoint is associated with the recommendation rating, on this client's data, for these customers. Not how much anyone assumes it should be.
Step three: forecast from the earlier reading
The transactional programme runs continuously and captures the customer at the moment of the interaction. The relationship survey captures the same judgement later, and less often.
Once the lag is known, the transactional reading becomes a leading indicator of the relationship reading, and the weights convert that into a forecast. What comes out is a predicted relationship score. The actual one still comes from the relationship survey when it reports; the prediction tells you what to expect before it does.
What is forecast is the recommendation rating, not the NPS score, and the distinction is not pedantic. NPS is percent promoters minus percent detractors: a banded calculation with thresholds at 6/7 and 8/9. A movement in the average rating does not convert into a fixed number of NPS points, because how much of it crosses a threshold depends on where your respondents were already sitting. The same one-point rise can move the score a lot or barely at all. Forecasting the rating keeps the model on a continuous measure and leaves the banding to be applied afterwards, on the client's own observed distribution.
How far ahead you can see. The full relationship reading can only be forecast as far out as your shortest touchpoint lag, because beyond that the short-lag inputs for that future wave have not happened yet. Longer-lag touchpoints give you early warning on their own contribution rather than on the whole number, and that is still worth having: it is the difference between finding out about a problem now and finding out about it in four quarters.
A worked example, constructed
The figures below are constructed to show the shape of the output. They do not describe any client.
A client runs three transactional programmes and a continuously fielded relationship survey, triggered on each customer's anniversary, carrying a rating question for each of the three touchpoints.
Touchpoint | Beta from regression | Measured lag |
|---|---|---|
Onboarding | 0.18 | 1 quarter |
Customer support | 0.41 | 2 quarters |
Claims | 0.27 | 4 quarters |
Customer support is associated with the recommendation rating more than twice as strongly as onboarding, per customer who passes through it. Whether it is worth more than twice as much to the headline number is a different question, and the caveat below answers it.
Claims carries a middling weight and the longest lag. A claims improvement made this quarter will not show up in the relationship reading for a year. And a claims deterioration this quarter has already locked in a decline four quarters out: it is visible in the transactional data today, and by the time the relationship wave reports it, nothing can be done about that wave.
How to read a beta, stated carefully. A beta of 0.41 means that, holding the other rated touchpoints constant, a customer who rates support one point higher rates the recommendation question 0.41 of a point higher. That is an association measured across customers, not a promise that raising support scores by a point will raise the aggregate by 0.41.
And a beta says nothing about how many customers are exposed. It is a per-customer intensity. As a planning heuristic rather than a calculation, what a touchpoint is worth to your average recommendation rating scales with its beta, the share of customers who pass through it, and how much you can realistically move it. The same association caveat applies to that heuristic as to the beta itself. A touchpoint with a beta of 0.18 that everyone experiences can be worth far more than one with a beta of 0.41 that a tenth of customers reach. Read the column as relative intensity and pair it with your own exposure data before deciding anything.
With that caveat attached, the table is a management document as much as a statistical one. It tells you where the leverage is, and when to expect the answer.
What this is actually for
The obvious reading is that once you can predict the relationship score from the transactional ones, you could stop running the relationship survey. That is technically true, and it is not what happens.
What the analysis buys is warning, not replacement.
A relationship score that has dipped is a question in a board meeting, and that is a bad question to be answering for the first time in the room. The dip is in the recommendation rating, the recommendation rating was heading down for weeks or months before the wave reported it, and a team running linkage already knew. They know which touchpoint carried it, roughly how much of the movement it explains, and whether it has already been addressed.
So the same meeting goes differently. Instead of a CX lead undertaking to look into it, there is an explanation ready before anyone asked for one, prepared weeks in advance rather than assembled overnight.
Companies keep running the relationship survey regardless, because it is the number that goes into the board pack, and being able to predict a number does not make anyone want it less. What changes is that the person presenting it is no longer surprised by it.
That is the honest value proposition here, and it is worth stating plainly rather than dressing up as prediction for its own sake. Preparedness is what clients actually buy.
What linkage analysis does not do
It is not a button in the platform. It is analysis alongside the licence: modelling work with time lags in it, done by data scientists on a client's own data.
Saying so plainly is a feature rather than an apology. Linking two instruments on different scales with different lags is a modelling problem, and any vendor claiming their platform does it automatically is simplifying something that does not simplify.
It does not produce a Numr benchmark for the lag. The lag is client-specific, and it varies by transaction process even within one client. Numr does not publish a lag figure, because the observation from one set of engagements may not hold across another industry or another client. The answer to "how long is the lag" is that it is measurable on your data, and that measuring it is the first step of the work.
It does not close the gap, and it is not meant to. It uses one reading to anticipate the other. The two instruments stay separate, because they are still measuring two different things.
What is deliberately left out of the model, and why
Earlier in this document, brand and pricing were named as things that sit in the relationship number and that no CX programme reaches. Both do influence the recommendation rating. Neither is put into the driver model, and that is a choice rather than an oversight.
Brand is excluded because it is itself a function of things outside the relationship, most of which have nothing to do with any touchpoint a client operates. Modelling it as a driver alongside onboarding and claims treats an aggregate of external forces as though it were an operational lever, which it is not.
Pricing is excluded because, in the programmes where Numr has tested it, it performs badly as a driver of the overall relationship. It reads as though it should matter and the data has not supported giving it a coefficient. Including it has degraded the model rather than improving it.
What that means in practice is that a linkage forecast is conditional on brand-level and price-level conditions holding. When it misses, that is worth knowing rather than embarrassing: a forecast built from journeys that fails to land is a signal that something moved which the journeys do not carry. Reading the residual is part of the analysis rather than a footnote to it.
A note on what NPS can and cannot carry
NPS is a useful shared language. It is not a complete measurement system, and the evidence for its supremacy has always been contested.
The metric was introduced by Frederick Reichheld in the Harvard Business Review in December 2003. Four years later, Timothy Keiningham, Bruce Cooil, Tor Wallin Andreassen and Lerzan Aksoy published a longitudinal examination in the Journal of Marketing, using Norwegian Customer Satisfaction Barometer data covering 21 firms and more than 15,500 interviews, and reported that they found "no support for the claim that Net Promoter is the 'single most reliable indicator of a company's ability to grow.'" That paper won the Marketing Science Institute's H. Paul Root Award in the same year.
More recently, Lance Bettencourt and Mark Houston, writing in the International Journal of Market Research in 2024, found that the reasons customers volunteer for their score are not reliably the things that actually move it. The two overlap, in their words, only to a "low to moderate" degree. That is an argument for measuring the drivers rather than reading the comments, which is the direction step two takes.
The practical reading is not to abandon NPS. It is to treat it as one instrument among several, keep the layers separate, and judge the programme by the changes it triggers rather than the number it reports.
Common mistakes
Blending the two into one score. A single averaged number destroys the only useful property either has, which is that they move at different times for different reasons.
Reading the direction of the gap as a diagnosis. Three mechanisms produce the same observable gap. Pointing at one of them without the data is guesswork.
Setting gap reduction as a target. It is the wrong target and it drives the wrong behaviour, which is usually to make the transactional survey more like the relationship survey.
Assuming the lag rather than measuring it. Every published lag figure describes somebody else's programme.
Changing several things at once during feed-through. If the relationship score moves two quarters after three simultaneous changes, you have learned nothing about which change did it.
Frequently asked questions
Why is our transactional NPS higher than our relationship NPS?
Both directions are normal, and the direction on its own is not a diagnosis. It could be feed-through lag, coverage of only some journeys, or something outside the journeys entirely such as pricing or brand. Telling those apart is a question for your own data.
What is the usual gap between transactional and relationship NPS?
There is not one, and there cannot be. The gap is a product of which touchpoints you measure, how fast each feeds through, and brand effects outside the programme. Anyone quoting a standard delta is describing one programme and calling it a rule.
How do we reduce the gap?
You should not set that as a goal. If the two converged you would have two surveys measuring the same thing and you would have lost the earlier signal. Watch each of them against itself instead.
We fixed the journey and the relationship score has not moved. What is wrong?
Most likely nothing. Touchpoint performance moves the relationship score with a lag, and each touchpoint has its own, so the relationship survey reports it later. The thing not to do is change something else in the meantime, because then you will not know which change did what.
Can you connect our transactional scores to our relationship score?
Yes, provided two conditions hold. Both surveys have to carry a rating question about the touchpoint, though not the same question and not on the same scale, and you need enough data points, normally at least a year. Given those, the lag for each touchpoint is measured from the two trends rather than assumed, each touchpoint is weighted by regressing its rating against the recommendation question inside the relationship survey, and the earlier transactional reading forecasts the recommendation rating ahead of the next wave. Numr calls this linkage analysis.
How long is the lag?
It is client-specific and it varies by transaction process. It is measurable on your data, and measuring it is the first step of the work rather than an assumption fed into it.
Does the model predict our NPS score?
It predicts the recommendation rating, which is the continuous 0 to 10 answer. NPS is a banded calculation on top of that, and a movement in the rating does not convert into a fixed number of score points, because how much of it crosses the promoter and detractor thresholds depends on where your customers already sat. The banding is applied afterwards, against your own distribution.
Do you model brand and pricing as drivers?
No, deliberately. Brand is a function of forces outside any touchpoint you operate, and pricing performs badly as a driver of the overall relationship. Both sit in the recommendation rating, and both are left out of the driver model, which means a forecast is conditional on those conditions holding.
Our relationship survey only runs once or twice a year. Can we still do this?
Not as it stands. The lag is found by sliding two series against each other, so the sparser one sets the resolution, and a survey fielded twice a year gives two points. The fix is not more surveying but different triggering: field the relationship survey continuously, on an anniversary or a renewal date, so it produces a rolling series while any individual customer is still asked once a year.
If linkage predicts the relationship score, can we stop running the relationship survey?
You could, and almost nobody does. The relationship survey is the number that goes into the board pack, and predicting it does not remove the want for it. What linkage changes is that nobody is surprised by it: a dip is understood, attributed and often already being fixed before the wave reports it.
Should we run both, or start with one?
Run the transactional layer first if you have to choose. It is the one that tells an operations team what to fix. The relationship layer is what tells an executive whether it worked, and it needs the transactional layer underneath it to be interpretable at all.