Two people in a meeting disagree about whether the thumbnails should have faces on them. Both have examples. Both are sincere. The argument has happened three times.

It will happen a fourth time, because nothing about the previous three produced a fact.

An experiment is how you end it: not by being more persuasive, but by agreeing in advance what would settle it and then going and finding out.

What separates an experiment from just publishing things

Four things. Miss any one and you've made content, not evidence.

  1. One variable changes. Everything else held as constant as you can manage.
  2. A benchmark frozen before you start. Written down, not recalculated afterwards.
  3. A sample size and a stop point decided in advance.
  4. A decision rule agreed before you see the result.

The fourth is the one that separates real experiments from expensive theatre. If you haven't agreed what result changes your behaviour, you will interpret whatever happens as supporting the position you already held. Everyone does this. Writing it down first is the only known defence.

Step 1: Turn the argument into a testable claim

"Faces work better" isn't testable. Too vague about what, for whom, measured how.

A testable version names the change, the metric and the direction:

"Thumbnails with a clearly-legible face will get a higher click-through rate than our current product-only thumbnails, on our long-form YouTube uploads."

Now it can be wrong, which is what makes it worth testing.

One variable. If you change the thumbnail and the title style in the same test, you learn nothing about either. This is by far the most common way content experiments get wasted. The temptation to improve two things at once is enormous and it destroys the result.

Step 2: Freeze the benchmark

Write down, before anything changes, what "normal" currently is.

Use the median, not the mean, over your last 20–30 comparable posts. One runaway drags an average far enough to make an ordinary result look like a win, which is exactly the error a benchmark exists to prevent.

Comparable means: same platform, same format, same rough era. A benchmark mixing Shorts and long-form is measuring the mix, not the content.

Step 3: Decide the sample size now

The question isn't "how long shall we run it". It's "how many pieces before I'd believe this".

There's no universal number, and anyone who gives you one is guessing about your variance. What's reliable is the shape of the problem:

  • Fewer than about five per arm tells you almost nothing. Content performance has a long tail; two freak results in five is an expected outcome, not a signal.
  • The noisier your channel, the more you need. Look at your existing spread. If your last thirty posts range from 4,000 to 90,000 views, a three-post test cannot detect anything smaller than an earthquake.
  • Decide it in advance and stop there. Stopping the moment the result looks good is how you guarantee a false positive: you're sampling until you like the answer.

Give the sample time to mature, too. A post's numbers at 48 hours and at three weeks are different numbers, and platforms distribute on very different curves. Compare like-aged posts or you're measuring age.

Step 4: Write the decision rule

Before you look. One sentence:

"If face thumbnails beat the frozen median CTR across eight uploads, we switch the default. If they don't, we stop having this argument and keep product-only."

Note what it commits you to in both directions. A rule that only says what happens if you're right isn't a rule.

Step 5: Run it, then read it honestly

When the results come in, ask three questions before drawing any conclusion.

Did anything else change? A platform change, a seasonal effect, a piece that got picked up somewhere. If your test period contains a confound, say so. A test run during a category-wide news event is measuring the news event.

How big is the difference relative to your normal spread? If your posts routinely vary by 40% and your test arm is 12% ahead, you have not found anything. This is the single most common misreading: small differences on small samples are noise wearing a result's clothing.

Is this [measured or reported](/learn/measured-vs-reported)? Did a system count it, or did someone summarise it? A dashboard screenshot with a hand-picked date range is a claim, not a measurement.

What you're allowed to conclude

Less than you'd like, and being disciplined about this is what makes the next experiment worth running.

You can conclude: this change, on this channel, in this period, moved this metric by roughly this much.

You cannot conclude: it will keep working, it'll work on another platform, it'll work for a different subject, or that it caused a downstream business outcome you didn't measure.

Audiences also habituate. A structure that wins in March may be worn out by August. Treat a result as a finding with an expiry date, not a rule.

The experiments actually worth running

Most teams have four or five recurring arguments. Those are your test list. Typically:

ArgumentThe one variable
Faces vs product on thumbnailsThumbnail style only
Long vs short formFormat, same topic
Hook structureOpening move, same content after
Posting cadenceFrequency, same content
Topic A vs topic BSubject, same format and treatment

Run one at a time. Two concurrent experiments on the same channel interfere with each other, and you'll be unable to attribute either result.

The uncomfortable part

A properly run experiment will often come back null: no detectable difference.

That is a real and valuable result. It means the variable you were arguing about doesn't matter much on your channel, and you can stop spending meetings on it and go look at something that does. Teams routinely treat a null as a failed experiment and quietly bury it, which is how the same argument comes back in six months.

Write nulls down. They're the cheapest thing you'll ever learn.

Where Acumin fits

Experiments is built around exactly this shape: you record what you changed, the benchmark is frozen at the point you start, and the result is measured against that frozen figure rather than against a moving average. Outcomes are read from real fetched data, not self-reported. That matters, because the thing you most want to fudge is an experiment that didn't go your way.

The Leaderboard ranks your own measured improvements, so the tests that actually moved something stay visible instead of being forgotten.

What none of it does is choose your variable or write your decision rule. Those are judgement, and they're the parts that make an experiment worth running.

How to use this tomorrow

Write down the argument your team has most often. Turn it into one testable sentence, name the metric, and write the decision rule (both branches).

That's fifteen minutes, and it's most of the work. The rest is just publishing things you were going to publish anyway.


Related: Measured vs reported is how to tell whether your result is a fact. Outliers, not virality is what to do with the data you already have, before running anything new.

Written by
Adam Murray
Founder, Acumin

Adam builds Acumin. He spends his days on the same two problems this library is about: working out what a piece of content is actually worth, and getting a brief through production without it turning into something else.

Want this done on your own channel?

Acumin reads your public content and your category and hands back what to make next. Every call comes with its confidence and the evidence it rests on. The first Snapshot is free.