How often should you check your AI visibility?
Monthly, against a question set you never edit. We measured the alternative on ourselves, and the run-to-run variation is the number most vendors do not publish.
Check monthly, on a frozen question set. Monthly is slow enough that a page you published has been re-crawled and reconsidered, and fast enough to notice a rival being absorbed into the answers. Weekly mostly records the engine rephrasing itself; daily records nothing you can act on at all.
That is easy to assert and we would rather show it. So we ran the same measurement on ourselves twice in a row, kept both records, and published the difference.
What run-to-run variation actually looks like
On 19 September 2026 we put eight buyer questions to Perplexity and analysed the answers. Then we did it again, about seventy seconds later, with nothing changed — same questions, same engine, same code. Both records are stored, and our own score page publishes them in full.
The headline measurements did not move at all:
| Measurement | First run | Second run |
|---|---|---|
| Visibility score | 0 | 0 |
| Answers naming Bungad | 0 of 8 | 0 of 8 |
| Mention rate | 0% | 0% |
| Citation share | 0.0% | 0.0% |
| Brands named, and in which answers | Identical across all eight answers: Profound (2), Peec AI (2), Otterly (2), AthenaHQ (1), Scrunch AI (1) | |
| Citations returned in total | 90 | 90 |
The sourcing underneath them did move. Both runs returned the same total of 90
citations and spread them differently: 74 distinct domains in run A, 75 in run B, and
the count for an individual domain drifted by one or two either way — the first run
credited blog.hubspot.com five times where the second credited it four.
The full side-by-side is on what Perplexity cites.
The rule that follows from it
Two runs seventy seconds apart cannot contain any real change in the world. So every difference between them is measurement noise, and that gives you a clean split:
- Treat the score, the mention rate and the brand tallies as measurements. They reproduced exactly across two independent runs of a non-deterministic engine.
- Treat an individual source count as an estimate. It moved when nothing had changed. Acting on a single domain gaining a citation is acting on noise.
- Treat the shape of the source list as a measurement. Which publishers and directories the engine reaches for held steady between the runs even as the counts shuffled. That is the level worth working on.
Ask your vendor for this number. Any tool selling you a trend line is subtracting run-to-run variation it has either measured or guessed. If it has measured it, it can tell you what it is. We publish ours because a trend without a noise floor is decoration.
How big a move has to be, and why your plan decides it
There is arithmetic here rather than judgement, and it is the most practical thing on this page. Being named carries half the score and rank among the brands named carries a fifth, so one answer flipping from ignoring you to naming you first moves the score by a fixed, calculable amount — and that amount depends only on how many questions you track.
| Questions tracked | One answer starts naming you first | What that means |
|---|---|---|
| 8 (our own audit) | Score moves about 9 points | Enough to be certain about a zero, too coarse to read a trend |
| 25 (Starter) | Score moves about 3 points | A move of five points or more is probably real |
| 100 (Growth) | Score moves about 1 point | Small moves become readable |
The size of your question set is your noise floor. That is the honest reason a larger tracked set costs more, and it is also the honest limit on our own published audit: eight questions is a sample big enough to establish a zero and far too small to read meaning into a single-point move later. Which is why we say so on the page that publishes it. Plan sizes and prices are on what AEO costs per month.
Why weekly is worse than monthly, not just more expensive
- Nothing you did has landed yet. A new page has to be crawled, indexed and then chosen as a source. That takes days to a couple of weeks. A third-party listing takes longer. A weekly chart of an unchanged web is a chart of the engine.
- The noise is bigger than the signal. On a small set, the move you are hoping to see is the same size as the move you get for free. You cannot tell them apart by looking harder.
- It makes you undo your own work. This is the real cost. A dip that is noise reads as a failure, and the response is to change the page that was actually working. Slower measurement protects good decisions.
When to check off-schedule anyway
Four cases, all of them verification rather than trend-watching:
- You published a page for a specific tracked question. Re-ask that question and check whether your page is now among the sources. You are looking for a citation, not a score move.
- An engine said something false about you. That is an incident. Fix the source and re-ask to confirm the correction propagated — when AI says something wrong about your brand has the procedure.
- A competitor launched. Worth one off-cycle run to see whether the answers have absorbed them, because that changes which questions are worth contesting.
- You added a question. Run it once to establish its starting point, and note the date you added it.
Freeze the question set. Really freeze it.
The discipline that makes a cadence worth having is not the interval, it is the question list. A score belongs to a set of questions as much as to a brand, so editing the set silently restarts the series while the chart carries on looking continuous. That is the most common way a visibility trend ends up meaningless.
Our own eight are fixed and published, word for word, precisely so next month's comparison means something. If you need to change yours, change it on purpose, write down the date, and read the trend from there. Choosing the questions to track covers how to pick a set you will not want to edit, and what counts as a good score covers how to read the number once you have a series.
What we do
Every paid brand gets one scheduled check a month, plus manual checks on top of it — four a month on Starter, twenty on Growth, a hundred on Scale — which exist for exactly the off-schedule verification cases above rather than for daily watching. Free accounts get manual checks only; the scheduled monthly run is what a subscription buys.
We hold ourselves to the same cadence in public. The same eight questions, once a month, published on our own score page whichever way the number moves.
Common questions
How often should you check your AI visibility?
Monthly, against a question set you do not change between runs. Monthly is slow enough that anything you published has been re-crawled and fast enough to catch a competitor moving. Weekly mostly measures the engine rephrasing itself, and a set of questions you keep editing produces a chart that cannot be read at all.
Does an answer engine give the same answer twice?
No, and we measured how much it moves. We put the same eight questions to Perplexity twice, about seventy seconds apart, and kept both records. The visibility score, the mention rate, the citation share and the tally of which brands were named in which answers came back identical. The sourcing underneath did not: the same total of 90 citations was spread over a different number of domains, and individual domains gained or lost a citation or two.
How big does a change have to be before it is real?
Bigger than one answer, and how much that is depends entirely on how many questions you track. On a set of eight, one answer that starts naming you first moves the score by about nine points on its own. On a set of twenty-five it moves it by about three, and on a hundred by about one. The size of your question set is your noise floor.
Should I check daily?
No. Nothing you can do about AI visibility has a daily cycle: a page has to be published, re-crawled and then chosen as a source, and a third-party listing takes longer still. A daily chart of a metric whose inputs move on their own is a chart of the engine, and reading it as progress leads to changing things that were working.
Can I change my tracked questions between checks?
You can, but the moment you do, the comparison with earlier runs stops being valid and you have started a new series. Add questions deliberately and rarely, note the date, and read the trend from that date forward. A score is a property of a question set as much as of a brand.
When is it worth checking off-schedule?
Four cases justify it. After you publish a page written to answer a specific tracked question. After you find an engine stating something false about you, because that is a correction to verify rather than a trend to watch. When a competitor launches and you want to know whether the answers have absorbed them. And when you first add a question, to establish where it starts.
Start the series
Enter your domain for the first run. Three questions on Perplexity, free, and the result is stored so the next one has something to be compared against.
3 questions on Perplexity, free. A domain checked in the last week is served from that stored run rather than asked again, and the daily free allowance resets at midnight UTC.