Two people bring you the same sentence: "Short-form is working better for us."
The first counted the views on your last thirty public posts. The second opened your platform analytics and compared watch-through. Same claim, same confidence in the voice, wildly different weight. And on a slide, in a meeting, they look identical.
That's the problem this article is about. Not "is the number right", but "what kind of number is it". Get that wrong and you'll make a large decision on evidence that was only ever good enough for a small one.
The ladder
There are four rungs. Each one sees something the one below it can't.
| Tier | Where it comes from | What it can settle |
|---|---|---|
| 1 | Public data: anything visible without logging in | What drew attention |
| 2 | Your own platform analytics | What held attention, and who it reached |
| 3 | Paid platform results | What earned its spend |
| 4 | Sales and CRM data | What moved the business |
The rungs aren't a maturity model you graduate through. They're four different instruments. A tier-1 read on a hundred posts can be far more reliable than a tier-2 read on four. What the tier tells you is the kind of question the evidence is equipped to answer.
Tier 1: public data
Views, likes, comments, upload dates, titles, thumbnails, and the same for everyone you compete with. No login, no permission, no cooperation from anyone.
What it can settle: what pulled attention. Which topics get clicked. Which formats get posted and which quietly got abandoned. How your output compares to your category in volume and in reception. It's also the only tier that works on other people: you will never see a competitor's analytics, and public data is the entire reason competitive research is possible at all.
What it can't: anything about the experience after the click. A video with a great title and a terrible middle looks identical to a great video at this tier. Public engagement also compounds with reach, so a post that was pushed hard by a platform outperforms a better post that wasn't, and tier 1 cannot tell the two apart.
Use it for: category reads, competitor work, and finding candidates worth a closer look. Not for judging your own creative quality.
Tier 2: your platform analytics
Impressions, reach, watch-through, retention curves, traffic sources, audience demographics. Yours, read-only, from the platform itself.
This is the biggest single jump on the ladder, because it's the first tier that separates distribution from content. Tier 1 shows you a post got 40,000 views. Tier 2 shows you whether those views came from search, from a recommendation, or from your own subscribers, and whether they stayed.
What it can settle: what held attention, where people left, which formats actually retain, and who your audience genuinely is rather than who you assumed. A retention cliff at 0:38 is the most actionable single fact in content, and it is invisible from the outside.
What it can't: whether any of it produced a customer. Retention is not intent. Your best-retaining video may be your least commercial one, and analytics will never tell you that.
Use it for: creative decisions. Format choices, structure, length, hooks.
Tier 3: paid results
Spend, impressions, click-through, cost per result, and the creative-level breakdown behind them.
Paid is a purpose-built experiment. You can put the same message behind two different hooks, hold the audience constant, and read the difference (something organic can almost never give you cleanly, because organic distribution is itself a variable you don't control).
What it can settle: which creative earns its money, at what cost, against which audience. It also, uniquely, tells you where organic and paid disagree. That disagreement is usually the most interesting thing on the page. Content that performs organically and dies when boosted was working because of context. Content that flops organically and flies when boosted was good, and simply never got shown.
What it can't: tell you much about the top of your funnel with any honesty. A tier-3 read on awareness content, judged on conversions, will kill your best awareness work.
Tier 4: sales and business data
Revenue, pipeline, attribution, cohort behaviour.
What it can settle: the only question that ultimately matters. Which content moved the business.
What it can't: be believed easily. Attribution across a multi-touch, multi-week, multi-device buying journey is genuinely hard, and most attribution models are opinions expressed as arithmetic. Tier 4 is the top rung, but it is not automatically the most trustworthy. It's the tier where the evidence is most valuable and the methodology most contestable.
The rule that actually matters
Match the tier to the size of the decision.
- Choosing a thumbnail? Tier 1 is fine.
- Choosing a format to commit a quarter to? You want tier 2.
- Choosing where to put a real budget? Tier 3, or you're guessing with money.
- Choosing whether to keep making content at all? That's a tier-4 question, and if you can't answer it, say so rather than dressing up a tier-2 answer.
The failure mode isn't using low-tier evidence. It's using low-tier evidence for a high-tier decision and not saying so.
Sample size is a separate axis
Tier is provenance. Sample size is reliability. They are not the same thing and confusing them is the second most common error here.
A tier-2 finding drawn from four posts is thin. A tier-1 finding drawn from three hundred is not. When you see a claim, you need both facts: where did this come from, and how much of it is there?
This is what a confidence band is for. It's the honest range around an estimate given how much evidence is behind it, and its real job is to make thin evidence visible, because a thin finding and a solid one produce the same-looking sentence otherwise.
How to use this tomorrow
Take the last three content decisions you made. For each one, write down two things: which tier the evidence came from, and how many observations were behind it.
Most people find at least one decision where the answer is "tier 1, about five posts, and we spent forty thousand on it".
That's the whole exercise. You don't need better evidence for every decision. You need to know which decisions are currently resting on evidence that can't hold them.
Related: Measured vs reported covers the other half of this problem: the difference between a number a system computed and a number someone told you. Outliers, not virality is about what to do with tier-1 data once you have it.