DELISORIA
Auf Deutsch
Portfolio01Magazine02Contact03 Auf Deutsch DE
Back to all articles

Magazine · · 3 min read

When a number means something — and when it is noise

A small website produces small numbers, and small numbers vary on their own. Read like large ones, they turn chance into a trend. How to tell the difference.

Eighteen more visits are not news

An analytics page shows 62 visits on a Monday and 44 the Monday before. That is 41 per cent more, and the figure sits there large and green. Almost every analytics tool shows exactly this — and almost every one leaves out the question that comes first: would the same difference have appeared if nothing had changed at all?

The smaller the number, the larger the share of it that is chance.

It would. Counts vary on their own, and the smaller they are, the more. At around a hundred counted events the usual variation is about ten; at ten thousand it is about a hundred — more in absolute terms, ten times less in proportion. That makes percentages the most dangerous way to present small quantities: 44 against 62 looks just as decisive as 4,400 against 6,200.

The square root as a rule of thumb

No statistics package is needed to estimate this. When you count independent events, the usual variation is of the order of the square root of their number. Comparing two periods with n events between them, a difference is only worth remarking on once it exceeds roughly two square roots of n.

A difference is only a difference once it exceeds what counts of that size vary by on their own.

For the 44 and the 62: 106 together, square root a little over 10, twice that is 21. The difference is 18 — below it. Nothing follows from these numbers. For 4,400 and 6,200: 10,600 together, square root 103, twice that 206, difference 1,800. That is movement.

This is the calculation behind every change in our own analytics. Under the arrow pointing up or down, in small type, it says ‘above the noise’ or ‘within the noise’ — and below 25 counted events it says nothing at all, because the rule of thumb no longer carries. An arrow without that note would be a claim.

The mean often describes no day at all

Thirty days, 197 visits, so 6.6 a day — that sounds like an ordinary day. Look into the days and three of them carried all 197 while 27 carried none. The mean here describes not one of the thirty. It is arithmetically correct and useless as information.

Where mean and median lie far apart, the mean describes nobody.

The median helps: it is the value half the days fall below, and a single spike does not lift it. When it sits far under the mean, that is the actual news — the traffic arrives in waves rather than as a stream. When it is zero, there is no baseline yet, and our analytics says exactly that instead of passing off a mean as an ordinary day.

The same caution applies to rates. A conversion rate from seven visits jumps between zero and fourteen per cent depending on whether one person happened to click. That is why we show rates over time by week and not by day: the jumping of a daily rate looks like a development and is not one.

Better to show nothing than something invented

The most uncomfortable decision in building an analytics page is leaving a panel empty. A table that computes shortfalls for twelve pages out of a single conversion looks like knowledge and is none. So we set thresholds everywhere: below five conversions in the period the impact table does not appear, below five distinct addresses no distribution curve, below seven days no distribution of daily values. In their place stands a sentence saying why.

An empty panel with a reason is more honest than a full one without.

That costs space and looks at first like a shortcoming. It is the opposite: a number with two decimal places looks like a fact even when it comes from three events. Whoever builds an analytics page decides, with every threshold, how much self-deception it permits.