What is anomaly detection?

Anomaly detection is finding the data points that do not fit the pattern of the rest. For business numbers, it is the difference between a sales figure that is merely different and one that is wrong. This explainer covers the kinds of anomaly, the common methods, and the part most tools get wrong.

Explainer 8 min read

Anomaly detection is the practice of identifying observations that differ significantly from what is expected, given the data around them. It is also called outlier detection or novelty detection, with small differences in meaning. In business, the observations are usually a metric over time, such as daily revenue, weekly spend or hourly orders, and the question is: is this reading unusual enough to deserve a person’s attention?

The key word is expected. An anomaly is never unusual in absolute terms; it is unusual compared with something. Most of the craft is in choosing that something well.

Three kinds of anomaly

Point anomalies

A single value far from everything else: a day with ten times the usual orders, a payment of $48,000 in a ledger of $200 transactions. These are the easiest to see and the easiest to detect.

Contextual anomalies

A value that is normal in one context and abnormal in another. $7,000 of revenue is an ordinary Tuesday and a bad Saturday. Heating costs that are normal in January are strange in July. Most real business anomalies are contextual, because business data has strong weekly, monthly and seasonal rhythms.

Collective anomalies and level shifts

A run of values that are each unremarkable but together are not: two weeks of revenue all slightly below normal, or a metric that steps down to a new level and stays there. No single day raises a flag. The pattern does. These are the anomalies that cost the most, because they are noticed last.

Common methods, and when each fits

MethodHow it worksGood forWeak at
Fixed thresholdAlert above or below a set valueHard limits, like a balance floorAnything with a rhythm
Z-scoreDistance from the mean, in standard deviationsStable, symmetric dataData with outliers, which distort both
Robust scoreDistance from the median, in median absolute deviationsMessy real business dataVery small histories
Seasonal baselineA separate normal for each weekday, month or hourContextual anomaliesNeeds several cycles of history
Changepoint testTests whether a series changed level, and whenLevel shifts and slow driftsOne-day spikes
Machine learningModels such as isolation forests learn many signals at onceMany related variablesExplaining why something was flagged

For a single business metric over time, the combination that works best in practice is a robust, seasonal baseline for day-to-day readings plus a changepoint test for level shifts. Machine learning methods earn their complexity when there are dozens of related signals, such as fraud across many transaction features, and struggle to say in plain words why a point was flagged.

A fixed threshold on six weeks of shop revenue: 6 alerts, every one an ordinary busy weekend, and the genuinely bad Wednesday missed.
The same data against a range learned per weekday: 1 alert, on the bad Wednesday.

Why false alarms matter more than people expect

A detector that flags too much is worse than it sounds, because people stop reading it. And the arithmetic of false alarms is unforgiving.

That is why a good detector needs more than a threshold. Three habits cut false alarms sharply: judging each reading in its own context, requiring a larger move for noisier metrics, and notifying when something changes rather than on every reading while an excursion continues.

How to tell if detection is working

  • Precision: of the things it flagged, how many were worth looking at?
  • Recall: of the real problems, how many did it flag?
  • Delay: how long after a problem started did it say so? A level shift found in three days is worth more than one found in three weeks.

The simplest honest test is a replay: run the detector over the last year of your own data and read what it would have flagged. If you would have ignored most of it, the sensitivity is wrong.

Pitfalls specific to business data

  • Missing data is not zero. A day with no export looks like a collapse in sales. A detector should say “no data”, not “anomaly”.
  • Holidays are real context. A closed day is expected to be empty.
  • Growth moves normal. A baseline that never updates will call steady growth an anomaly forever. See rolling baselines.
  • Ratios hide their parts. A margin can hold while revenue and costs both fall. Watch the parts as well as the ratio.

In short

  • An anomaly is a reading that does not fit what is expected in its context.
  • Most business anomalies are contextual or collective, not single spikes.
  • Robust, seasonal baselines beat fixed thresholds for metrics with a rhythm.
  • False alarms, not missed detections, are what usually kill an alerting system.

Let a gauge do the watching.

Free to start: two gauges, no card needed.