Running your first experiment

The return leg of the loop. Record what you changed, let the benchmark freeze, and read the result against a number that can't move underneath you.

Adam Murray6 August 20267 min read
On this page

Everything else in Acumin runs in one direction: data → decision → work. Experiments is the return leg: the part that tells you whether the decision was any good.

Without it you're producing more confidently. With it you're producing more accurately, which is a different thing.

What the surface does

Three things, and the middle one is the point:

  1. You record what you changed (the one variable).
  2. The benchmark freezes at the moment you start.
  3. The result is measured against that frozen figure, from real fetched data.

Why "frozen" is the load-bearing word

A benchmark that updates as new results arrive quietly absorbs whatever you just did. Your new content gets compared against a baseline that now contains your new content, and the comparison becomes circular.

You end up measuring your work against itself, which reliably produces a small, encouraging, meaningless improvement.

So the snapshot is taken at the start and left alone. That's not a technical detail. It's the entire reason the result means anything.

Before you set one up

Three decisions the surface can't make for you, covered properly in running a content experiment that settles the argument:

One variable. If you change the thumbnail and the title style, you learn nothing about either. This is the most common way content experiments get wasted, and the temptation to improve two things at once is enormous.

A sample size, decided in advance. Fewer than about five per arm tells you almost nothing: content has a long tail, and two freak results in five is expected behaviour rather than a signal. Decide the number before you start and stop there; stopping the moment it looks good is how you guarantee a false positive.

A decision rule, written before you look. "If it beats the frozen median across eight uploads we switch the default. If it doesn't, we stop having this argument." Note that it commits you in both directions. A rule that only says what happens if you're right isn't a rule.

Reading the result

The surface gives you the observed delta against the frozen benchmark. Three questions before you conclude anything:

Did anything else change? A platform shift, a seasonal effect, a piece that got picked up somewhere. A test run during a category-wide news event is measuring the news event.

How big is the difference relative to your normal spread? If your posts routinely vary by 40% and your test arm is 12% ahead, you have not found anything. Small differences on small samples are noise wearing a result's clothing.

Is this [measured or reported](/learn/measured-vs-reported)? Here it's measured: outcomes are read from real fetched data, not self-reported. That matters more than it sounds, because the thing you most want to fudge is an experiment that didn't go your way.

Nulls are results

A properly run experiment will often come back with no detectable difference.

That is a real and valuable outcome. It means the variable you were arguing about doesn't matter much on your channel, and you can stop spending meetings on it. Teams routinely treat a null as a failed experiment and quietly bury it, which is exactly how the same argument returns in six months.

Write them down. They're the cheapest thing you'll ever learn.

Where results surface afterwards

The [Weekly Digest](/learn/making-sense-of-the-weekly-digest) has a "what's measurable now" section: tests that have accumulated enough data to read. That's the prompt to go and look, and it's the main defence against setting up an experiment and forgetting it.

The [Leaderboard](/app/leaderboard) ranks your own measured improvements, so tests that actually moved something stay visible rather than being forgotten. It ranks your own work only: nothing cross-user, nothing forecast.

What to test first

Your team's most recurring argument. Everyone has three or four:

  • Faces versus product on thumbnails
  • Long versus short form, same topic
  • Hook structure, same content after
  • Posting cadence

Run one at a time. Two concurrent experiments on the same channel interfere, and you'll be unable to attribute either result.

How to use this tomorrow

Write down the argument your team has most often, turn it into one testable sentence, and set the experiment up today, before you make the change.

The setup is fifteen minutes. The frozen baseline is the thing you can't go back for.


Related: Running a content experiment that settles the argument is the method in full. From signal to brief to shoot is the outbound half of the loop.

Written by
Adam Murray
Founder, Acumin

Adam builds Acumin. He spends his days on the same two problems this library is about: working out what a piece of content is actually worth, and getting a brief through production without it turning into something else.

Want this done on your own channel?

Acumin reads your public content and your category and hands back what to make next. Every call comes with its confidence and the evidence it rests on. The first Snapshot is free.