Evidence

Does AEO actually work?

Parts of it are mechanically certain and the rest is unproven, including by us. No company in this category — ours included — has published a controlled result, so here is what is actually known, what would count as proof, and the test you can run yourself in thirty days.

Published 19 September 2026 · 6 minute read

We sell this service, so treat the page accordingly and check every claim on it. We have tried to write the version that survives being checked.

What is mechanically certain

A short list, and everything on it follows from how retrieval-based answering works rather than from anybody’s case study:

If someone is doing only these four things for you, they are doing work that is certain to be necessary. That is not the same as certain to be sufficient.

What is plausible but unproven

What nobody has published, and why that matters

We looked for a controlled study — a fixed question set, a treated group and an untreated control, measured before and after, with the raw data available. We could not find one from any vendor in this category, and we have not run one either. What circulates instead is single-brand before-and-afters with no control, which cannot distinguish the intervention from the category simply getting more coverage over the same period.

The specific thing we will not tell you is how long it takes. We have one month of our own data and no client history long enough to answer honestly. Every duration figure we have traced in this category came from somebody’s estimate. Ask us again after our third monthly audit, when the answer can be a chart rather than a guess.

What would count as proof

  1. A fixed question set, published. If the prompts change between measurements there is no comparison, and a report that does not print its prompt list cannot be trusted to have kept them stable.
  2. A baseline with the raw answers stored, not just a score. A score you cannot audit is a number somebody chose.
  3. Repeat runs at each point. Engines are non-deterministic: our two runs a minute apart moved individual domain counts by one or two while the headline held. One run is not a measurement.
  4. A control. Questions you deliberately do not work on. Without them you cannot separate your effort from the whole category moving.
  5. Publication either way. A method that only gets published when it worked is marketing.

The test you can run yourself, in thirty days, for nothing

  1. Day 1. Write ten questions a buyer asks before they know you. Split them into six you will work on and four you will deliberately leave alone. The four are your control.
  2. Day 1. Ask all ten, twice, in a logged-out session. Record for each: were you named, in what position, and the full list of cited sources. Save the raw text. This is your baseline and its value depends entirely on being honest now.
  3. Days 2 to 20. For the six, write one page per question. Title it as the question, answer it in the first forty words, put the specifics in text. The method is on writing a page an engine will quote. Change nothing about the other four.
  4. Day 30. Ask all ten again, twice, the same way. Compare the six against the four.
  5. Read it honestly. If the six moved and the four did not, you have weak but real evidence. If both moved, the category moved. If neither moved, thirty days was too short or the pages are not answering the question as directly as the pages that are being cited.

This costs nothing but your time and it is a better experiment than anything currently published, which is an indictment of the category rather than a compliment to the test.

The experiment we are running on ourselves

Our audited visibility score on 19 September 2026 was 0: named in 0 of 8 answers, 0 of 90 citations. We published it in full before we had anything good to report, and we are re-running the same eight questions on the same engine every month.

The falsification condition, stated in advance: if bungad.com is not appearing in the citation list for at least one of the eight questions by the third monthly run, the thesis that publishing question-shaped pages moves AI visibility is wrong for our case, and the answer is third-party sources — directories, review platforms and publications — rather than more pages of our own. We will publish that result as prominently as we would publish a good one.

Naming the condition before the data arrives is the only way a published result means anything. It is on our score page and in our content plan, and you can hold us to it.

So should you buy it?

On the evidence available today: do the four mechanically-certain things yourself first, because they are free and their necessity is not in doubt. Measure before you start, so you can tell later. Then decide about paying anyone — including us — with a baseline in hand rather than on a promise, using tool, agency or in-house to work out which shape of purchase you are even making. If your site is under twenty pages, spend the money on writing instead; that recommendation is on AEO for a small business on a budget and it costs us sales.

Common questions

Is there proof that AEO works?

Not of the kind that should convince you. We could not find a controlled study from any vendor in this category, with a fixed question set, a treated group, an untreated control and published raw data, and we have not run one either. What exists is single-brand before-and-afters with no control, which cannot separate the intervention from the whole category getting more coverage over the same period.

What part of AEO is definitely worth doing?

Four things that follow from how retrieval works rather than from anyone's case study: a page that does not exist cannot be cited, a page that cannot be fetched cannot be cited, a fact that is not in text cannot be quoted, and the engines do retrieve from the live web. Those are certain to be necessary. They are not thereby certain to be sufficient.

How long does AEO take to show results?

We will not give you a number. We have one month of our own data and no client history long enough to answer honestly, and every duration figure we have traced in this category came from somebody's estimate rather than a measurement. Ask us again after our third monthly audit, when the answer can be a chart.

How can I test whether it works for my own site?

Split ten buyer questions into six you will work on and four you will deliberately leave alone as a control. Measure all ten twice on day one and store the raw answers. Write one page per question for the six over the next three weeks. Measure all ten again on day thirty. If the six moved and the four did not, that is weak but real evidence. It costs nothing but time.

Take the before measurement

A test needs a baseline. Three of your buyers’ questions on Perplexity, free, with the date and every source recorded, so you have something to compare against later.

3 questions on Perplexity, free. A domain checked in the last week is served from that stored run rather than asked again, and the daily free allowance resets at midnight UTC.