FoundationsFor brands

Reading your numbers without fooling yourself

The statistical and cognitive traps that make content data lie to you: small samples, survivorship, seasonality, cherry-picked windows and false cause.

Adam Murray4 September 20268 min read
On this page

You picked the right metric. It came from the platform, so it was counted rather than asserted. And you can still draw a completely wrong conclusion from it, because the number is honest and your reading of it isn't.

That gap is the subject here. There's a separate gap worth keeping distinct: measured vs reported is about where a number came from: whether something counted it or someone claimed it. That's the attribution gap. This piece assumes the number is genuinely measured and asks a different question: even so, are you thinking about it correctly? These are the thinking traps: the ways a real number leads a reasonable person somewhere false. They're worth knowing by name, because each has a specific fix.

Small samples are mostly noise

The single most common error in content analysis is treating a handful of posts as a trend.

Content performance has a long tail. A small number of posts do far better than the rest, and which post that is has a large element of chance: the algorithm's mood that hour, what else was in the feed, who happened to be online. When you have five posts, one lucky one can swing the average enough to invent a pattern that isn't there. Change one post and the "trend" reverses.

The fix is to refuse to conclude from too few. As a rough working rule, fewer than about fifteen posts in one platform-and-format isn't a baseline yet. It's a rumour. And use the median, not the mean: one runaway post drags a mean upward until everything else looks like failure, which is exactly backwards when the runaway is the thing you wanted to study. Outliers, not virality leans on this deliberately: a single outlier is an anecdote; the finding only counts when the pattern repeats across enough posts to rule out luck.

You only see the posts that shipped

Survivorship bias is the trap of studying the winners and never the ones that didn't make it.

It shows up twice in content. First, in your own work: you analyse what you published, but the ideas that got killed in the draft stage (the ones that would have flopped, and the ones that would have flown but scared someone) never enter the data. Second, and more dangerously, in what you copy from others. The formats that "always work" are the ones visible enough to notice. You never see the thousand accounts that ran the same format into silence, because silence doesn't trend. So the format looks far more reliable than it is.

The fix is to actively look for the missing cases. When a pattern seems to work, ask: where are the failures of this same pattern, and would I even be able to see them? And apply the discipline from outliers, not virality: check whether your losers share the trait you credited your winners with. If your top five all open with a question and so do your bottom five, the question-opening is a habit, not a signal. A trait only counts if it's present in the wins and absent from the losses.

Seasonality moves the floor

A number went up. Did your content improve, or did the calendar?

Audiences, categories and platforms all have rhythms. Traffic sags in a holiday week and swells when your category is in the news. B2B goes quiet on Fridays and in December. If you compare this month to last month without accounting for the season, you'll credit the calendar to your creative or, worse, blame a good post for launching into a naturally quiet window.

The fix is to compare like periods and know your own rhythm. Where you can, hold the season constant: this month against the same month last year, or this Tuesday's post against your Tuesday median, rather than against Saturday. And when the category itself was in the news the week a post spiked, write that on the post before you conclude anything about its format. Timing masquerading as content is one of the easiest false lessons to bank.

A window can be chosen to say anything

Give someone a long enough time series and they can find a start and end point that proves nearly any story. Start the chart at your worst week and today looks like a triumph. Start it at your best and you're in decline. Same data, opposite conclusions, both technically true.

This is cherry-picking, and the reason it's so seductive is that it usually isn't dishonest. It's motivated. You choose the window that matches what you already believe, without noticing you chose. The screenshot crop is the physical form of the same act: the pixels are real, the framing is a claim, as measured vs reported puts it.

The fix is to decide the window before you look at the result. Fix the comparison period in advance ("we'll judge this quarter against last quarter") and hold to it even when a kinder window is available. If you must show a shorter window, show the longer one beside it. And be suspicious of any chart, including your own, whose start date is doing a lot of the persuading.

Correlation is not cause

Two things moved together. That is not evidence that one caused the other, and content data is full of pairs that rise together for a third reason entirely.

You started posting more and follows went up, but you also started posting more because a product launch was drawing attention, and the launch drove both. Video length correlates with watch time, but longer videos are also made by your more experienced people on your stronger topics. The confound is usually invisible precisely because it's the thing driving everything.

The fix isn't a statistics degree; it's a habit of asking one question before you act: what else changed at the same time? If you can name a plausible third factor, you don't yet have a cause. You have a hypothesis. The way to promote a hypothesis to a finding is to change one thing on purpose and watch, which is what a deliberate content experiment is for. Until then, hold the causal claim loosely and say "associated with", not "drove".

Where Acumin fits, and where it can't

Some of this can be built into the tooling and some can't. Confidence bands exist precisely because of the small-sample trap: when there isn't enough data to say something firmly, the honest output is a wide band or "not enough to say", not a crisp number that pretends to a precision the data doesn't hold. A Snapshot reads your own baseline from real public data, which is what lets you compare a post to your norm instead of to a cherry-picked week.

But no tool can save you from the last trap. A model can compute the correlation; it cannot know what else changed in your business the week the line moved. That judgement is yours, and it's the reason the interpretation is deliberately left to a human rather than generated as a confident sentence. The tool's job is to give you a number it won't inflate; your job is to not fool yourself with it.

How to use this today

Take the most confident conclusion you currently hold about your content: the "X always works for us" you'd say in a meeting without checking. Run it through the three questions in the callout above: how many posts, what else changed, was the window chosen after the fact.

If it survives all three, trust it more than you did. If it doesn't, you've just avoided building next quarter on a coincidence, which, more often than anyone likes to admit, is the more valuable outcome. Then take the metric you settled on in the metrics that matter and hold it to the same test before it drives a single decision.

Written by
Adam Murray
Founder, Acumin

Adam builds Acumin. He spends his days on the same two problems this library is about: working out what a piece of content is actually worth, and getting a brief through production without it turning into something else.

Want this done on your own channel?

Acumin reads your public content and your category and hands back what to make next. Every call comes with its confidence and the evidence it rests on. The first Snapshot is free.