Paste your analytics into a chat window and ask what's working. You'll get an answer. It will be well-structured, appropriately hedged, and contain figures.
Some of those figures will be correct. Some will be the most plausible number for that sentence, produced by the same process that produced the words around them.
You cannot tell which is which by looking. That's the problem, and it's architectural rather than a matter of the model being good or bad.
Why it happens
A language model generates the most likely next token given everything before it. That's the whole mechanism, and it's why it's so good at prose.
When the preceding text is "engagement rose by", the most likely continuation is a plausible percentage. The model has no separate faculty that checks whether a computation was performed. Producing "34%" and producing the word "significantly" are the same operation.
This isn't a flaw to be trained away. It's what the thing is. A system whose job is producing plausible continuations will produce plausible numbers, and a plausible number is indistinguishable from a computed one once it's in a sentence.
The four things it genuinely can't do
Compute reliably over your data. Even with real numbers in the prompt, arithmetic performed as text generation is unreliable in ways that don't announce themselves.
Know its own uncertainty about a fact. It can express hedging language, because hedging language is a pattern. That's not the same as knowing whether a specific figure is grounded, and the two are easily confused because they look identical in output.
Distinguish your data from its training data. Ask about your channel's performance and you may get an answer partly shaped by general patterns about channels. Fluently blended, not flagged.
Tell you what it wasn't given. If you paste three months of data and ask about the year, you'll typically get an answer about the year.
What it's genuinely excellent at
This isn't an argument against using models. Handed verified numbers, they're very good at the things that are actually hard:
- Explaining what a figure means in context
- Structuring an argument from a set of established facts
- Spotting the question you should have asked
- Drafting the memo once the analysis exists
- Arguing with your interpretation, which is underrated
The division that works: let it reason and write; don't let it originate facts.
The architecture that fixes it
The fix isn't better prompting. Prompting reduces the frequency of invented numbers and cannot eliminate the category, because the mechanism producing them is the mechanism producing everything else.
The fix is structural: compute figures in code, from real data, and give the model only those figures to write about. The model composes prose around numbers it is not permitted to originate.
That's a real constraint with real costs. Sometimes the honest output is "there isn't enough data to say", where an unconstrained model would produce something satisfying. You lose fluency at the edges and you gain the ability to act on the output.
What to do if you're not using such a system
Most people are pasting data into a general chat window. That's fine, with three habits:
Ask where every number came from. "Which of these figures did you calculate, and from what?" A model will often reveal that it was estimating. Not reliable, but cheap.
Verify anything you'd act on. Any figure going into a decision, a deck or a client conversation gets checked against the source. Every time.
Prefer questions with no numeric answer. "What might explain this pattern?" is a good use. "What's my engagement rate?" is a bad one: you can compute that exactly, and asking a model to is inviting the failure mode for no benefit.
The scale problem
One point that gets missed. This used to be a small issue because the volume of analysis was small: a human wrote a monthly report, and if a number was wrong someone had a chance of noticing.
The volume of generated analysis is now enormous. Every team, every week, in every tool. The rate of unverifiable numbers entering business decisions has gone up by a large multiple, and the mechanisms for catching them have not improved at all.
That's the real reason measured vs reported matters more now than three years ago. The distinction is the same; the volume of reported-numbers-wearing-measured-clothing is not.
Where Acumin fits
This is the constraint the product is built around, and it's the reason for a design that would otherwise look over-engineered.
Numbers come from code. Computed by the pipeline from real fetched data: public content, connected analytics. Words come from the model. The interpretation, the phrasing, the argument. Confidence comes from the data: how much there was, and how consistent.
The model composes sentences around figures it cannot originate. Which is why it will sometimes tell you the evidence is thin instead of producing a number that would fit nicely.
Ask Acumin carries a specific version of this: any figure in a sentence about your own brand's performance has to trace to a real tool result, or it doesn't survive. General numbers (dates, specs, public facts) aren't touched. The guard is aimed at exactly one failure mode.
None of that makes the interpretation right. It makes the interpretation checkable, which is the only property that matters when you're deciding something.
How to use this tomorrow
Take the last AI-generated analysis you received about your own performance and pick out every number in it.
For each one, answer: could I reproduce this from source data in under five minutes? The ones where the answer is no are the ones to stop repeating.
Related: Why honest numbers matter more now is the argument for why this is getting more important. Measured vs reported is the underlying distinction.