Ask a few. Act for everyone.

Behavior shows where a customer stopped, never why. How the reasons a few customers give extend to the ones who never answered.

Amitayu Basu CEO and Co-founder
9 min read
A key hanging from a lock in a dark wooden door, lit from one side.

Your stack already knows the customer is stuck, and it knows where. What it cannot say is how to get them moving again; that was never its job.

The objective has not changed: solve the customer's problem, and the customer moves forward. What that takes is knowing the problem, and only the customer can tell you what it is.

TL;DR

  • Your systems already see where a customer is stuck; behavior never says why.
  • A survey supplies the why, but only for the minority who answer.
  • PXI extends those reasons to silent customers as calibrated estimates, via classification on behavior. Where no theme stands out, it skips the customer.
  • A theme that stands out is paired with its experiment-approved intervention. Proof comes from randomized group experiments.
  • It sits alongside the core platform and the CRM. Nothing is displaced.

A question with no owner

Your core banking platform runs accounts and transactions, and your CRM holds the record and the workflow; both do exactly what they were bought to do. What neither was ever asked is what is blocking a customer, or what would get one moving again. In a standard stack that question has no owner, not because anything fell short, but because it was never given to anyone. It is an absence, not a failure.

An action chosen without the problem

Any decision about a customer draws on many inputs, but two of them do the work here. The first is behavior: what the customer did and where they stopped, all of it already in your stack. The second is stated reasons: what the customer said was wrong, in their own words. The method described below needs both, yet only the second carries a reason.

That asymmetry has a cost. A reason inferred from a click is a guess until it is checked against reasons customers actually stated, and behavior alone contains nothing to check it against. So the action can be well timed, well targeted, and still the wrong move: the right customer, reached at the right moment, offered a fix for a problem they do not have. The customer stays stuck.

Call it a next best action program or anything else, the decision runs on its inputs, and what would unstick this customer is simply not among them.

Only a conversation gives you a reason

If you want to know why customers get stuck, someone has to tell you. That is because the vocabulary of reasons is either stated by customers or guessed on their behalf; no third place exists for it to come from.

A survey is simply that conversation run at scale: short, close to the moment, in the customer's own words. Its output is not a score, and nobody here is asking you to care about scores; it is a set of reasons, tied to the customers who gave them.

The objection you are already forming

Only a minority answer, so stated reasons cover a sample, not the base. If the reasons stop at the respondents, you have a research finding, not an operating input.

And the sample is not neutral either. In digital collection people tend to answer when something has an emotion attached, so the strongly satisfied reply, the strongly dissatisfied reply, and the indifferent stay silent. Respondents therefore differ from the silent in composition, not just in count.

Both objections are correct, and they are where PXI starts rather than stops.

Ask a few. Act for everyone.

Start with the reasons stated by the sample. The theme set comes from there, derived from customer language and then governed and labeled, so in one specific sense nothing is invented: no reason category exists that a customer did not state.

The per-customer assignment is a different kind of thing, and we name it plainly: an estimate. Supervised classification, trained and calibrated on survey-supplied labels, produces the probability that a governed theme applies, given behavior. That is neither a statement by the customer nor a guess, but an estimate with a known error behavior, which is a third kind of thing.

Prediction itself covers everybody: every customer is scored against every theme, producing a spread of probabilities rather than one label. The discipline sits one step later, in the decision to nudge. When one theme stands out, the customer qualifies for its nudge, and when two stand out, they can qualify for two. But when the spread is flat and nothing stands out, we skip that customer, because we would rather say nothing than nudge at a problem we cannot identify.

The projection rests on an assumption, and we would rather state it than have you find it. It assumes that, conditioned on behavior, silent customers carry a similar spread of reasons to the customers who spoke. Nothing above discharges that assumption. What makes it reasonable rather than idle is this: silence alone cannot tell an indifferent customer from a blocked one, and behavior carries material information about which is which. So behavior is a non-arbitrary thing to condition on, and the assumption stands as an assumption, in the open.

So the method needs both inputs: the where from your record, the why from the customers who spoke. The survey stays the source of the reasons; the product adds the calibrated projection and the discipline of when to act on it.

For the data science reader, the shape of it is worth stating. The input is two things you already have or can gather: the behavioral record and the survey responses. The output, per customer, is a calibrated probability against every theme. The validation is the kind you would ask for: calibration is measured, and so is the classifier's error behavior.

Even so, that output is a ranking input rather than an answer. It says who is more likely to be stuck and on what; what to do about it is a separate question, answered separately.

That makes the title literal: you ask a few, you predict for everyone, and you act wherever the prediction is clear enough to act on.

Where the claim stops

It is worth being precise about what is predicted. We predict, per customer, that they are likely to stall and on what theme; we do not predict, per customer, whether a given intervention will work on them. That second judgment belongs to the experiment, not to the model.

This is a narrower claim than it could be, but it is the defensible one, because a model can rank likelihoods while causal credibility comes only from randomization. The experiment that supplies it is described next.

The proof loop

It runs in three steps.

  1. Test. Randomized treatment and control measures the causal effect of each intervention policy. The readout is standard: difference in outcome rates, treatment effect, minimum detectable effect, proportion tests, confidence intervals, power.
  2. Publish. The intervention that earned its lift becomes the approved intervention. Targeting then runs on the theme that stands out for a customer, through that approved mapping.
  3. Track. Group lift is measured on what was published. That measurement is the proof shown.

The honest formulation is short: we target where a theme stands out, and we prove by group experiment. Causal credibility comes from randomization, not from a model; a model says who is likely to stall and on what theme, but only an experiment says whether the intervention changed the outcome. A before-and-after chart is not proof.

Where it sits

The core platform keeps running the bank, whether that is Fiserv, Jack Henry, or Q2, and the CRM keeps holding the record and the workflow. PXI is an addition alongside both.

It reads what you already send and adds the reason layer. What it hands back is a theme and an action, delivered into the tools you already run, such as Salesforce or Braze. The action itself comes from the experiment-approved mapping, not from the model.

Nothing is displaced and nothing is re-platformed; there is no migration and no system of record to move. The platform page states the framing plainly: production without rip-and-replace. The banking-specific view of the same architecture is on the retail banking page.

The integration question has a short answer. What does it displace? Nothing. What does it need? What you already send, plus the survey it runs.

One more point, because it matters here: PXI does not instrument your customer journeys. The behavioral record is yours, from your own systems, sent on your terms.

On contacting customers

In many of our programs, and most consistently in our European ones, permission to follow up is captured inside the survey itself, and agreement gates whether an alert is generated at all. Consent is therefore a survey-design decision, made before any alert exists. It also means the contactable population is smaller than the stalled population, which we treat as a property of the design, not a flaw.

Not a churn product

Churn prediction asks who is likely to leave, and it is a real discipline; we have written about how customer churn prediction works. PXI asks a different question: who started something and did not finish it? An application was opened and abandoned; a product was adopted and never funded. That revenue is not lost; it is stuck.

For the customers who answered, the reason is known; for the rest it is estimated, and nudged wherever the estimate is clear enough to act on.

Frequently asked questions

Does this replace our CRM or our decisioning tools?

No. The core platform keeps running the bank, and the CRM keeps the record and the workflow. PXI adds a reason layer alongside them and hands actions back into the tools you already operate.

Where does the behavioral data come from?

From your own systems. PXI does not instrument journeys, track hesitation, or capture screen-level abandonment. It reads the behavioral record you already hold and choose to send.

Do you predict which message will convert a specific individual?

No. We predict, per customer, the likelihood of a stall and its theme. Which intervention earns the right to be published is decided by randomized experiment, and lift is measured at group level.

What if survey response rates are low?

The design assumes it. The respondents supply the reasons, and classification carries them, as calibrated estimates, to every customer. Where no theme stands out for a customer, no nudge goes out.

How is timing decided?

By controlled policy experimentation. Timing policies are tested the same way messages are: randomized, measured, and published only when they earn it.

Amitayu Basu CEO and Co-founder

25 years in customer experience, helping global brands listen. Numr is what he built when listening stopped being the hard part.

Share

Stop measuring the stall. Start recovering the revenue.

Watch the 2-minute demo