Reading the number

What is a good AI visibility score?

There is no published benchmark, so any industry average you are quoted is unsourced. What we can give you is the band table, the arithmetic ceiling, and the one comparison that means anything.

Updated 22 September 2026 · 6 minute read

Nobody can honestly tell you the average AI visibility score, because no census of them exists. On our own scale, under twenty-five is the rarely surfaced band and anything in the seventies or above is consistently surfaced. But the useful comparison is not against a benchmark or against a competitor — it is against the same brand’s previous run, on the same questions, on the same engine.

This page is the interpretation layer. How we measure AI visibility gives the formula and the worked arithmetic; this one answers the question people actually ask after their first check, which is is this bad?

Why no benchmark exists

A visibility score is not a temperature. It is a weighted summary of three measurements taken over a list of questions somebody chose, on whichever engines the vendor covers, in one category. Change any of the three and the number changes without anything about the brand changing.

So an “industry average AI visibility score” would be an average over incompatible measurements. When you see one quoted, ask which question sets it pooled and which engines were in it. We have never seen that answered, which is why we do not quote one either.

A grid of dots on a raised panel, a column for each run and a row for each tracked question, a filled dot where the answer named the brand, with the most recent column picked out in gold

The four bands, and what they are for

These are the labels our software prints beside a score. They exist to stop a bare number reading as a verdict, and they are deliberately coarse — a band is the most a single figure can honestly support.

ScoreBandWhat it describes
0–24Rarely surfacedThe engine almost never reaches for you when asked.
25–49Occasionally surfacedYou turn up, unpredictably, on some of your questions.
50–74Regularly surfacedYou are a normal part of the shortlist in your category.
75–100Consistently surfacedYou are hard to leave out, and usually cited as well.

There is a fifth state, and it is the one that matters most for honesty: not measured. A run that asked nothing at all is not a zero, and our banding function refuses to label it — bandOf returns unmeasured rather than rarely surfaced, because rarely surfaced is a finding and a failed run is not one.

The ceiling is lower than the scale suggests

Three inputs are weighted: how often you are named carries half, how much of the cited sourcing is yours carries three tenths, and where you come among the brands named carries the remaining fifth. A brand named first in every one of eight answers, with not a single citation pointing at its own site, scores 70.

That is worth sitting with, because it changes what a “good” number looks like. To go beyond the low seventies you have to stop being merely recommended and start being the source the recommendation is built on. A perfect 100 requires every citation in every answer to point at you, which we have never seen and do not expect to.

Practical reading: the seventies are the top of the realistic scale, not a near miss. If a vendor shows you a client sitting at 95, ask to see the citation share underneath it — either it is remarkable, or the questions were written to name the brand.

Zero is a finding. Blank is not.

Our own audited score is 0: eight buyer questions put to Perplexity, Bungad named in 0 of 8 answers, mention rate 0%, citation share 0.0%, no average position because we never appeared. The whole run is published at our own score.

That zero is a measurement, and it is not the same thing as a zero that means “we could not tell”. There are three genuinely different states a citation share of nought can describe, and collapsing them is how a report ends up asserting something it never observed:

Our software distinguishes these three rather than printing one sentence for all of them, because “it never cited your site” implies it cited somebody else, and on a run that cited nobody that sentence is an invention. Ours is the middle case: the engine cited 90 sources across those eight answers and not one of them was bungad.com.

The scores we could publish and will not

We hold the stored answers from our own audit, so we could compute a visibility score for every rival named in them and publish a league table. We do not, and the reason is the same one that makes benchmarks meaningless.

Those eight questions were chosen to measure us. Scoring another company against a list written for our own category, with no citation data gathered on their behalf and no say in the wording, would produce a number that looks like their measurement and is not. What we do publish is the raw observation, which needs no modelling: across all eight answers, identical in both stored runs, Profound was named in two, Peec AI in two, Otterly in two, AthenaHQ in one and Scrunch AI in one. Bungad in none.

Read that honestly and it is the most encouraging fact in our own audit. The best-performing brand in the category appeared in a quarter of the answers. A score in the twenties can be category-leading here. Nobody has consolidated this yet, and the comparison of the tools says the same thing from the buyer’s side.

When a high score is the symptom

The fastest way to a flattering number is to track questions containing your own brand name. Ask an engine what reviewers say about Acme and it will say Acme. The mention is handed over by the question, it cannot be lost, so it cannot move — and a metric that cannot move is not a measurement.

We treat that as a correctness rule rather than a guideline. A question naming the brand is still asked, the answer is still stored and shown, because what an engine says about you is genuinely worth reading. It is excluded from every figure, and the exclusion is printed next to it in the words the customer reads. Choosing questions that can be lost is the whole skill, and choosing the questions to track goes through it.

What to compare against instead

  1. Your own previous run, question set frozen. The only comparison with no confound in it. Change the questions and you have started a new series.
  2. The brands named in the same answers as you. Inside a single run, on identical questions, that comparison is fair — it is the same measurement applied to everyone in the answer. This is what citation share is for.
  3. The gap between named and cited. Being mentioned and being the source are different achievements, and the distance between your mention rate and your citation share tells you which half of the work is outstanding.
  4. Your worst single question. More actionable than any average. There is a page to write, and the question tells you what it is about.

If your own number came back low, the diagnostic order is in why your brand is not in AI answers, and how often to re-check it is in how often to check your AI visibility.

Common questions

What is the average AI visibility score?

Nobody knows, and anyone who quotes you a figure is quoting something they cannot source. There is no census of AI visibility scores, because there is no shared scale: every vendor weights its own inputs differently, over a question set the customer chose, on whichever engines that vendor happens to cover. An average across those is an average of different measurements.

Is a score of zero a broken measurement?

Not necessarily, and the two cases have to be told apart before the number means anything. A run that asked nothing is unmeasured and our software refuses to band it at all. A run that asked its questions, got answers, and found the brand in none of them is a finding. Our own audit is the second kind: named in 0 of 8 answers, score 0.

Can a brand score 100?

Only by being named first in every answer and being the source behind every single citation in all of them, which does not happen. A brand named first in all eight of our audit questions, with no citation pointing at its own site, scores 70. That is the realistic top of the scale, so treat the seventies as excellent rather than as a near miss.

Why can I not compare my score with a competitor’s score?

Because a score belongs to a question set as much as to a brand. Change the questions and the same brand scores differently, and no two brands track the same set. Comparing scores across brands compares the question lists. What is comparable between brands, inside one run, is who was named in the same answers.

What makes a score too high to trust?

A tracked question that contains your own brand name. The engine is told the name and hands it back, so the mention is guaranteed by the question rather than earned, and it can never be lost. We still ask such questions and keep the answers, because what an engine says about you is worth knowing, but they are excluded from every figure and the reason is printed beside them.

How much does a score have to move before it means something?

More than the engine’s own run-to-run variation, which you have to measure before you can subtract it. On a set of eight questions one answer changing its mind moves the mention rate by a whole eighth, so single-point moves on a small set are noise. Watch which questions changed rather than the headline number.

Get a number to interpret

Enter your domain and we run the measurement described above — the score, its band, who was named instead of you, and the sources the engine used. Perplexity, because it is the only engine we hold a key for.

3 questions on Perplexity, free. A domain checked in the last week is served from that stored run rather than asked again, and the daily free allowance resets at midnight UTC.