Sign in

Blog · Wallet share and penetration

Validate inferred wallets against known customer spend

Evaluate inferred category-spend estimates on customers withheld from the estimation process. Measure absolute error, bias and segment differences before expanding their use.

The short answerBacktesting an inferred wallet means producing the estimate without access to a customer's known category spend, then comparing it with that independently held-out evidence. Report dollar error, percentage error and directional bias by relevant segment. Keep the estimator and test population frozen so the exercise measures performance on unseen customers rather than reproducing the data used to build the norm.

Share-of-wallet estimate backtesting asks whether inferred customer spend is close enough to independently known spend for the decision you want to make. A neatly reconciled sales ledger cannot answer that question: it validates the numerator, while the uncertain denominator still needs its own evidence.

The norm-construction guide owns how an estimate is built. This guide evaluates an already specified estimator without changing it to fit the test answers.

Collect comparable declarations

Choose customers with dated category-spend declarations, tender totals or another reviewed source. Match the category, entity level, currency and period to the intended estimate. A group-wide annual figure is not a valid test for one site's quarterly wallet.

Store observed wallet, evidence date, reference period, source, category and review status. Preserve uncertainty in declarations: a rounded estimate supplied during a sales conversation is weaker evidence than a documented spend schedule. Do not present the observed value as perfect ground truth when it is not.

Hold out entire related customer groups where shared information could leak between subsidiaries. If the future deployment is in a different segment, show how many genuinely comparable customers the test includes.

Freeze the estimator before opening answers

Save its version, reference customers, size inputs, scaling rule and any minimum-group rule. Produce predictions from information that would have been available at the decision date. Do not use a customer's newly declared total to revise its size classification and then claim an accurate blind estimate.

Hyndman and Athanasopoulos explain why accuracy on unused data differs from fit to construction data. Their time-series cross-validation discussion also supports keeping future information out of historical tests. The customer-wallet procedure here is an application of those evaluation principles, not a prescribed forecasting model.

Use interpretable error measures

Choose the error sign explicitly:

Error = predicted wallet − observed wallet

Mean absolute error = sum of absolute errors / tested customers

Weighted absolute percentage error = sum of absolute errors / sum of observed wallets

Aggregate signed bias = sum of signed errors / sum of observed wallets

Calculate percentage errors only for positive observed wallets. A zero or unknown denominator is a separate exception, not a zero-error record. Aggregate signed bias can hide large offsetting errors, so report it alongside absolute error.

A synthetic held-out test

All amounts are USD thousands for matching annual categories. These four observations are invented.

Account Observed wallet Blind estimate Signed error Absolute error
A 100 110 +10 10
B 200 180 −20 20
C 400 480 +80 80
D 300 330 +30 30
Total 1,000 1,100 +100 140

Mean absolute error is $140,000 / 4 = $35,000. Weighted absolute percentage error is 14%. Aggregate signed bias is +10%, indicating overestimation in this sample. Individual absolute percentage errors are 10%, 10%, 20% and 10%; their unweighted average is 12.5%, a different statistic.

These results describe four test cases. They do not establish a universal acceptable error threshold or an industry benchmark.

Test the ranking decision as well

A moderate dollar error can still reverse two near-tied account priorities. Compare the gap ranking based on the blind estimates with the ranking based on reviewed declarations. Record which accounts crossed the team's pursuit cutoff and how large the difference was.

Use estimate sensitivity to show whether those priorities survive plausible denominator changes. A broad error distribution is a reason to widen scenarios or restrict use, rather than supply more decimal places.

Decide what evidence the result supports

Break out performance by segment, size band, source age and estimation method. Report missing cells and their sample counts. Do not claim a segment passed because a pooled result across unrelated customers looks acceptable.

Agree the intended use with the reviewer: exploratory account ranking, a customer conversation or a materially consequential plan may require different evidence. Retain failed cases and revise the method on a new development set; reserve fresh cases for the next test.

Bring the declaration sample, estimator version and error table to Covirage to discuss a bounded analysis. A backtest informs where inference is useful and where the evidence coverage should remain explicitly unknown.

Questions people ask

Can the same customers define the norm and test it?

They can describe fit to the construction data, but they do not provide an independent test. Keep test customers and related entities out of the reference group, or use a documented cross-validation process.

Does a small backtest prove accuracy for the whole book?

No. Customers willing to declare spend may differ from the unknown population. Report sample size, segment coverage, period alignment and selection limitations before applying the result elsewhere.