Skip to content

Evidence

Does AEO work?

There is one widely cited study, a great deal of vendor reporting, and very little in between. That is a thin evidence base for a category this loud, and it is worth knowing exactly how thin before spending against it.

The short answer

Partly, and less certainly than it is sold. The strongest published evidence is a 2023 benchmark study reporting visibility gains of up to forty percent from specific content changes on the generative engines of that time. Everything else in public is either vendor case studies without controls or anecdote. Directionally the advice holds. The size of the effect on your site is unknown until you measure it.

What ranked on 18 September 2026: A search for the academic term puts the arXiv paper first and fifth, Wikipedia fourth and Google's own guidance third. The commercial phrasing of the same topic returns no primary evidence at all. The two SERPs describe the same subject with completely different standards of proof.

What the one real study found

Aggarwal and colleagues framed the problem as optimising a source document for an engine you cannot inspect, then tested edits against a benchmark of queries.

  • Adding citations, quotations and statistics to a document raised its visibility in generated answers most reliably
  • Keyword stuffing, the tactic most visible in current AEO advice, performed poorly
  • The headline figure of up to forty percent is a benchmark result on the engines of 2023, not a client outcome and not a promise about today's models

Why vendor evidence is weak

This is not an accusation of dishonesty. It is a structural problem with measuring a moving target.

  • No control group. The same site cannot be run both ways at once, so improvement and model update are indistinguishable
  • Sampling noise. Ask the same question twice in a day and the answer may cite different sources, so small differences are not real
  • Selection. Case studies published are the ones that worked, which tells you the ceiling rather than the average
  • No shared definition of a citation, so two vendors can report different numbers for the same month and both be right

How to test it on your own site

A defensible in-house test is cheap, and it beats every published average because it is about you.

  • Fix a question set of twenty to forty prompts your buyers would plausibly type, and never edit it once sampling starts
  • Sample every engine that matters weekly, recording mention, position in the answer and how you are described
  • Change one class of thing at a time, for example restructuring twenty pages into self-contained answers, and hold the rest still
  • Give it six to eight weeks, and discard any week in which a model update was announced
  • Judge on the trend across runs, never on a single sample

What to do with this

  1. 01Treat the forty percent figure as evidence that the direction is right, not as a forecast. Anyone quoting it as a promised outcome has not read the conditions attached to it.
  2. 02Do the cheap parts regardless: self-contained answers, accurate structured data, correct third-party descriptions. They are durable and they help ordinary search too.
  3. 03Require any vendor to state how they would know their work caused a change. The answer tells you more than their case studies do.

Questions

Is there proof that AEO increases revenue?
Not in public, in any controlled form. There is a benchmark study showing content changes can raise the chance of being cited, and vendor case studies without controls. Revenue attribution from AI answers is harder still, because most engines send no referrer that identifies the answer as the source.
How long does it take to see results?
Page-level changes can appear in engine answers within days once the page is recrawled, but a trend takes six to eight weeks of sampling to distinguish from noise. Anyone promising a timeline shorter than their own sampling interval is describing hope.
What is the cheapest thing that actually helps?
Rewriting your most important pages so the first paragraph answers the question by itself, with the entity named rather than implied. It costs a day, it survives model changes, and it improves the page for human readers too.