Blog · Wallet share and penetration
Evaluate inferred category-spend estimates on customers withheld from the estimation process. Measure absolute error, bias and segment differences before expanding their use.
Share-of-wallet estimate backtesting asks whether inferred customer spend is close enough to independently known spend for the decision you want to make. A neatly reconciled sales ledger cannot answer that question: it validates the numerator, while the uncertain denominator still needs its own evidence.
The norm-construction guide owns how an estimate is built. This guide evaluates an already specified estimator without changing it to fit the test answers.
Choose customers with dated category-spend declarations, tender totals or another reviewed source. Match the category, entity level, currency and period to the intended estimate. A group-wide annual figure is not a valid test for one site's quarterly wallet.
Store observed wallet, evidence date, reference period, source, category and review status. Preserve uncertainty in declarations: a rounded estimate supplied during a sales conversation is weaker evidence than a documented spend schedule. Do not present the observed value as perfect ground truth when it is not.
Hold out entire related customer groups where shared information could leak between subsidiaries. If the future deployment is in a different segment, show how many genuinely comparable customers the test includes.
Save its version, reference customers, size inputs, scaling rule and any minimum-group rule. Produce predictions from information that would have been available at the decision date. Do not use a customer's newly declared total to revise its size classification and then claim an accurate blind estimate.
Hyndman and Athanasopoulos explain why accuracy on unused data differs from fit to construction data. Their time-series cross-validation discussion also supports keeping future information out of historical tests. The customer-wallet procedure here is an application of those evaluation principles, not a prescribed forecasting model.
Choose the error sign explicitly:
Error = predicted wallet − observed wallet
Mean absolute error = sum of absolute errors / tested customers
Weighted absolute percentage error = sum of absolute errors / sum of observed wallets
Aggregate signed bias = sum of signed errors / sum of observed wallets
Calculate percentage errors only for positive observed wallets. A zero or unknown denominator is a separate exception, not a zero-error record. Aggregate signed bias can hide large offsetting errors, so report it alongside absolute error.
All amounts are USD thousands for matching annual categories. These four observations are invented.
| Account | Observed wallet | Blind estimate | Signed error | Absolute error |
|---|---|---|---|---|
| A | 100 | 110 | +10 | 10 |
| B | 200 | 180 | −20 | 20 |
| C | 400 | 480 | +80 | 80 |
| D | 300 | 330 | +30 | 30 |
| Total | 1,000 | 1,100 | +100 | 140 |
Mean absolute error is $140,000 / 4 = $35,000. Weighted absolute percentage error is 14%. Aggregate signed bias is +10%, indicating overestimation in this sample. Individual absolute percentage errors are 10%, 10%, 20% and 10%; their unweighted average is 12.5%, a different statistic.
These results describe four test cases. They do not establish a universal acceptable error threshold or an industry benchmark.
A moderate dollar error can still reverse two near-tied account priorities. Compare the gap ranking based on the blind estimates with the ranking based on reviewed declarations. Record which accounts crossed the team's pursuit cutoff and how large the difference was.
Use estimate sensitivity to show whether those priorities survive plausible denominator changes. A broad error distribution is a reason to widen scenarios or restrict use, rather than supply more decimal places.
Break out performance by segment, size band, source age and estimation method. Report missing cells and their sample counts. Do not claim a segment passed because a pooled result across unrelated customers looks acceptable.
Agree the intended use with the reviewer: exploratory account ranking, a customer conversation or a materially consequential plan may require different evidence. Retain failed cases and revise the method on a new development set; reserve fresh cases for the next test.
Bring the declaration sample, estimator version and error table to Covirage to discuss a bounded analysis. A backtest informs where inference is useful and where the evidence coverage should remain explicitly unknown.
They can describe fit to the construction data, but they do not provide an independent test. Keep test customers and related entities out of the reference group, or use a documented cross-validation process.
No. Customers willing to declare spend may differ from the unknown population. Report sample size, segment coverage, period alignment and selection limitations before applying the result elsewhere.