Sign in

Blog · Data quality and reconciliation

Outliers: the one-off order that breaks the norm, and what to do with it

How a single large order distorts every measure built on it, the norm it inflates, the run rate it doubles, the concentration it spikes, the share of wallet above 100 percent it produces, how outliers are detected from the account's own history and the segment's distribution, the rule that they are flagged and shown rather than removed, the medians that make norms robust to them, and the two kinds of outlier that are not errors at all.

The short answerOne order ten times an account's usual size doubles its run rate, puts its share of wallet above 100 percent, spikes concentration, and if the account is in the norm's population, raises the norm for every similar account. Outliers are detected against the account's own history, an order beyond a stated multiple of its median, and against the segment's distribution. They are flagged and shown, never removed, because some are real: a project, a stock build, a new site. Norms use medians so one order cannot move them, and run rates use windows long enough to hold several orders. The report shows the measure with and without the flagged order, and the reader decides.

A mid-sized account places one order for a plant opening, ten times its usual size. The next month its share of wallet reads 180 percent, its run rate has doubled, its segment's norm has risen, and it is the company's fourth-largest customer. Every one of those is arithmetic on a real order and every one misleads. This guide sets out what an outlier does, how it is detected, the flag-and-show rule, and the two kinds that are not errors.

What one order does

Measure Effect
Run rate, three-month window Doubles; annualised figure is fantasy
Share of wallet Above 100 percent; the wallet estimate looks wrong
Concentration The account jumps into the top ten for a quarter
Norm, if computed as a mean Rises for every similar account; every gap widens
Dormancy, next year The account's cadence and typical size are distorted
Forecast bias The rep who forecast it was right once and wrong on the trend

Detection

Two tests per order, stated:

Own-history test: order value > k × the account's median order, trailing year, with at least n orders Segment test: order value > the stated percentile of the segment's order distribution

Either flags. The flag carries the reason.

The rule: flag and show, never remove

Where the order appears Treatment
Ledger and identity Included; the identity holds
Concentration Included, flagged on the account
Run rate Shown with and without; the without is the default for ranking
Share of wallet Shown with and without; the wallet estimate is not revised on one order
Norm Unaffected: norms are medians
Forecast Excluded from the trend; included in the actual

Medians make norms robust

A norm computed as a mean moves with one order. A median does not. Every norm on this site is a median or a percentile for that reason, and a mean is used only where the report says so and why.

The two kinds that are real

Kind Looks like Treatment
Project or opening order One order, one product family, then back to cadence One-off flag; excluded from run rate; kept in the ledger
Step change A large order, then a new higher cadence that holds Not an outlier after two more periods; the account moved band

The second is the customer growing, and the report re-reads it as a step change once the new level repeats.

A worked flag

Account Median order This order Multiple Segment percentile Flag Reason logged
4471 $9,000 $92,000 10.2 99.8th Yes Plant opening, per rep
2207 $40,000 $58,000 1.5 91st No
Account 4471 With order Without
Run rate, annualised $410,000 $112,000
Share of wallet 180% 49%
Rank by revenue 4 38

The reader sees both, and the reason.

Where it goes wrong

Removed. The identity breaks and a real order vanishes.

Not flagged. Share of wallet at 180 percent; the wallet method blamed.

Norms as means. One order moves every similar account's gap.

Step change flagged forever. The customer grew and is treated as an anomaly for a year.

Every upload, flagged and shown both ways

Mapped once, the ledger's history produces the two tests, the flags with reasons, and every affected measure with and without the flagged orders. Covirage builds this from the exports as they are. The metrics governance solution describes the setup, and the run rate guide covers the one-off order as one of the four ways run rate misleads.

Questions people ask

Why not remove outliers?

Because the order happened, the revenue is real and the ledger has it. Removing it breaks the identity and hides a fact. Flagging it and showing the measure both ways keeps the identity and gives the reader the choice, with the reason on the line.

How is an outlier detected?

Two tests, stated: the order is beyond a multiple, say five, of the account's own median order over the trailing year; and it is beyond a percentile of the segment's order distribution. Either flags it. An account with fewer than a stated number of orders uses the segment test only.

Which outliers are not errors?

A project order, a customer building stock ahead of a price change, a new site's opening order. They are real, they will not repeat, and the report treats them as one-off: excluded from the run rate window with a note, kept in the ledger and the concentration figure, and flagged on the account for the next period's comparison.