Skip to content

Why Most AEO Case Studies Prove Nothing: How to Read Them

Most AEO case studies skip baselines, controls and attribution. Use this six-question test to read them, with real examples and SparkToro data on variability.

28 Sept 20268 min read
  • Critique
A marketing team meeting around a table, illustrating Why Most AEO Case Studies Prove Nothing: How to Read Them

Most published AEO case studies do not prove that the tactic caused the result, because they rarely disclose baselines, tracking methodology, a control group or what else changed during the window. That does not mean AEO does not work. It means a "600% citation uplift" is a claim to interrogate, not a benchmark to copy. As of 30 September 2026, the fix is a short checklist you can apply to any case study in ten minutes.

Key Takeaways

  • A percentage lift without the raw before-and-after numbers is close to meaningless; going from 2 to 8 citations and from 200 to 800 are both "300%".
  • AI answers are highly variable, so a single before-and-after snapshot can be noise. SparkToro's research found repeated identical prompts almost never returned the same brand list.
  • Most public case studies lack control groups, disclosed attribution methods or independent verification.
  • Ask six questions: baseline, tool, window, confounders, business outcome, reproducibility.
  • Treat a missing answer as an answer. If the author will not share the number, assume it would not help their case.
  • Your own case study should be held to the same standard, including mine.

The news hook: AEO proof is suddenly everywhere

Search "AEO case study" and you will find dozens of pages promising proof of ROI. HubSpot's roundup of answer engine optimization case studies alone lists agency claims such as trials rising from 575 to 3,500+ a month in seven weeks, a "600% citation uplift", 10% of organic traffic from LLMs, and one law-firm example attributing $2.34M of revenue to AI discovery. I went through that page line by line, and the pattern is consistent: the execution is described in detail, the measurement is not.

I am not calling anyone dishonest. Agencies are allowed to summarise. But a summary is not evidence, and buyers are making budget decisions on these summaries.

Why AEO is unusually hard to prove

Three properties make answer-engine measurement noisier than classic SEO.

1. The outputs are non-deterministic

In research published by SparkToro (Rand Fishkin and Patrick O'Donnell, January 2026), 600 volunteers ran 2,961 prompt runs across ChatGPT, Claude and Google's AI Overviews/AI Mode using 12 prompts. Their finding: there was less than a 1 in 100 chance of getting the same brand list twice, and less than 1 in 1,000 for the same order. It was not peer-reviewed and used normal user settings, which the authors say was deliberate. Still, the implication is direct: one screenshot of "we now appear" is weak evidence.

The same research offers a constructive point. Visibility percentage across many prompts, run multiple times, was a reasonable metric, while "ranking position" in AI answers was not.

2. Everything changes at once

A typical AEO project involves schema fixes, new articles, digital PR and community posting at the same time. If citations rise, which change did it? And what did the model provider change that month? See the Reddit citation story in my post on that topic: a platform-side change can swing a whole category with no action from any brand.

3. Tools disagree with each other

Each visibility tracker samples different prompts, locations and models. A gain in one dashboard may not exist in another. Which brings us to the first thing to ask about any study.

A marketer marking up a printed case study with a red pen at a desk
Read case studies like a reviewer, not a buyer: pen in hand, looking for what is missing.

The six-question test

This is adapted from the criteria I use with clients, and it overlaps with the checklist published by W3Solved on reading AI visibility case studies.

  1. What were the raw baseline numbers? Not just the percentage.
  2. Which tool measured it, and is the method inspectable? Proprietary and unexplained is a flag.
  3. What was the window? Month-by-month is better than start and end points.
  4. What else changed? Migrations, algorithm updates, PR spikes, paid campaigns.
  5. Did it reach a business outcome? Leads, demos, revenue, not only mentions.
  6. Could someone reproduce the claim from what is disclosed?

Score a study one point per clear answer. Anything under four is a story, not proof.

Worked example: scoring four public claims

Using the claims as HubSpot's page reported them (I have not independently verified any underlying client data):

Claim (as reported)Baseline disclosed?Method disclosed?Confounders?Score
Trials 575 to 3,500+ in 7 weeks, 600% citation upliftTrials yes, citations noNoSite fixes, 66 articles, Reddit work all at onceLow
63% brand citation rate on awareness prompts (Apollo.io)PartlyNames tool (AirOps)Community plus contentMedium
10% of organic traffic from LLMs (3 months)NoNo attribution modelUnknownLow
$2.34M revenue from AI discovery (6 months)NoNo CRM or attribution windowUnknownLow

Notice that the strongest entry only scores "medium" and still lacks a control. That is the norm, not the exception.

The most common tricks (mostly accidental)

Percentages on tiny bases

Going from 2 to 8 monthly citations is a genuine 300% increase and a trivially small change in visibility. The percentage is true; the impression is false.

Cherry-picked prompts

If the prompts were chosen after seeing which ones improved, the result is survivor bias. Ask when the prompt list was fixed.

Visibility mistaken for value

Being mentioned is not being chosen, and being chosen is not revenue. Studies that stop at citation share should say so.

Attribution by assumption

AI-referred traffic is notoriously hard to attribute; many visits arrive with no referrer. Claims of "27% of AI sessions converted" need a described method.

What good evidence looks like

I would take seriously a study that: pre-registers a prompt set of at least dozens of prompts; runs each several times per week; reports visibility percentage with ranges; includes a comparison group of similar pages or products that were not changed; and reports pipeline outcomes against a documented attribution rule. Few agencies can afford this for every client, and that is fine, as long as they label weaker work honestly as "observations".

Where my own work fits

I work in organic growth, and the case study I reference most is Masai School. There I helped grow Instagram from 26K to 117K followers and LinkedIn from 50K to 200K, which you can read about on Masai School's site. Apply your own test: that is social growth, not AEO, and I cannot claim it proves an AEO method. It shows sustained organic work; it does not isolate one tactic. That gap between "result" and "proof of cause" is exactly the point of this post.

Line chart on a laptop screen showing a rising trend with no axis labels
A rising line with unlabelled axes is a picture, not a measurement.

How to run your own honest AEO test

You can do this without a big budget.

  1. Pick 30 to 50 prompts your buyers really use and freeze the list.
  2. Record current visibility by running each prompt at least three times across ChatGPT, Google AI Mode and one other engine. Store raw outputs.
  3. Choose a treatment set of pages and a similar untouched control set.
  4. Change only the treatment set for 8 to 12 weeks.
  5. Re-measure the same way, and report both sets.
  6. Log every external event: model updates, competitor launches, your own PR.

If the treatment set gains and the control does not, you have something. If both rise, the tide did it.

What to do when a vendor sends you a case study

Reply with the six questions. A serious vendor will answer some and admit what they cannot. SparkToro's Rand Fishkin advises asking AI tracking tools to demonstrate their methodology before you buy; the same principle applies to case studies. Vague answers are informative.

Caveats on this post

I did not have access to the underlying client data of any case study mentioned, so I assess disclosure, not truth. Vendor claims are attributed as reported. Nothing here says AEO is ineffective; it says proof is thinner than the marketing suggests.

FAQ

Do AEO case studies ever prove causation?

Rarely. A controlled test with a comparison group, a fixed prompt set and documented attribution comes closest. Most public examples are observational reports of a bundle of changes, so they show association at best.

Why do AI answers vary so much between runs?

Language models sample their outputs and often retrieve different sources per query. SparkToro's January 2026 research found identical brand lists appeared less than 1% of the time across repeated runs, so single snapshots are unreliable.

What is a good sample size for tracking AI visibility?

Nobody has established a standard. SparkToro's authors left open how many runs are needed. In practice I use dozens of prompts, each run several times, and report a visibility percentage with a range.

Is a big percentage lift always misleading?

Not always, but it is incomplete. Always ask for the raw numbers. A 300% rise from 2 to 8 citations is real yet tiny; the same rise from 200 to 800 is meaningful.

Should I trust vendor-published case studies?

Treat them as hypotheses. Vendors have a commercial reason to publish wins, so check methodology, confounders and whether outcomes reached revenue or only mentions.

How can I measure AI referral traffic?

Use analytics referrer data where available, add UTM parameters to links you control, and ask new leads how they found you. Expect undercounting, since many AI visits arrive without a referrer.

What should a control group look like?

Similar pages or products in the same category, left unchanged during the test. Without one you cannot separate your work from platform changes or seasonality.

Does this mean AEO is a scam?

No. Structuring content clearly and earning mentions are sensible practices. The problem is overclaiming: selling certainty where the evidence shows probability.

How long should an AEO test run?

At least 8 to 12 weeks in my experience, with weekly measurement, so you can see trend rather than a single spike. Longer if the model providers ship major updates in the window.

Want an evidence-first approach to organic growth?

I would rather show you what I measured and what I could not. With 4+ years of marketing experience across edtech and startup clients, I build SEO, AEO and content programmes that report honestly, including the limits. Have a look at my work and reach me through the contact form at https://younusfardeen.in.