How to

Choosing the questions to track

The ones a buyer asks before they have heard of you, phrased the way they would type them, and fixed for the whole series. Everything else — brand-name prompts, keyword fragments, a list you edit each month — produces a number that moves for reasons you cannot explain.

Updated 19 September 2026 · 6 minute read

This is the decision everything else rests on. The question set is the denominator of every metric you will report, so a bad set does not produce a slightly wrong number — it produces a number that is not about anything.

The method, in six steps

  1. Start from the sales call, not from a keyword tool. Write down what people actually ask you before they buy. Keyword tools return search phrases; answer engines are asked questions, and the two are not the same shape.
  2. Drop every question that contains your brand name. Someone who already knows your name is not choosing between you and anyone. Brand-name prompts flatter the report and measure nothing.
  3. Cover all five question types. Category, comparison, price, best-for-a-situation, and problem-shaped. A set that is all one type measures one narrow thing and will move all at once or not at all.
  4. Phrase each one the way a person types it. Full sentences, contractions and all. Do not compress a question into a keyword, because a shorter prompt returns a broader answer and a broader answer names more brands.
  5. Fix the list and write it down with today's date. Changing the questions between runs changes the denominator and destroys the comparison. Version the list; do not edit it in place.
  6. Record what the set cannot see. Every question you left out is a blind spot you have chosen. Write those down too, so the next round can close them deliberately rather than by accident.

The rest of this page is why each of those is the rule, and what goes wrong when it is broken.

The rule that matters most

Track questions a buyer asks before they have heard of you. A prompt containing your brand name measures whether an assistant can look you up, which it almost always can, for a question almost nobody asks. It will make your report look good and tell you nothing.

The test: would someone type this if they had never encountered your company? If not, it does not belong in the set. There is a place for brand-name prompts — checking whether an assistant describes you accurately — but that is a different job, covered on when AI says something wrong about your brand, and it should not be mixed into a visibility score.

The five types to cover

TypeShapeWhat it tells you
Category “What is X and who provides it?” Whether you exist in the engine’s map of the field at all.
Comparison “Alternatives to [competitor]” Whether you are in the consideration set. The highest-intent type, and usually the easiest to win.
Price “How much does X cost per month?” Whether your prices are in plain text where they can be quoted.
Best-for-a-situation “Best X for a small team on a budget” Whether you are matched to a segment, rather than just to the category.
Problem-shaped “How do I stop Y happening?” Whether you reach people who have not yet framed their problem as your category. The hardest and often the most valuable.

A set that is all one type moves as a block. Five types means that when the number changes you can say which kind of buyer changed their mind about you.

How many

Fewer than tool vendors suggest, and more than one. Roughly:

We track eight for ourselves. Our own plans allow 25 on Starter and 100 on Growth, and we would rather you used 25 well than 100 badly.

Phrasing: write it the way people type it

Instead ofTrackWhy
“AEO tools” “Best AEO tools for tracking brand visibility in ChatGPT and Perplexity” A two-word prompt returns a broad answer naming many brands. The longer question is what a buyer actually types, and it is winnable.
“AEO pricing” “How much does answer engine optimization cost per month?” The question form triggers a retrieval that matches question-shaped pages.
“Profound competitors” “Alternatives to Profound for AI search visibility tracking” Matches how the query is typed, and adds the qualifier that decides which alternatives are relevant.

One caution against over-tuning: a question you phrase to suit your own page will return your own page. Phrase it as a stranger would, then check whether you win.

Fix the list, then leave it alone

Every metric is computed over the set. Change the questions and the denominator changes, so month two is not comparable with month one, and the trend — the only number that ever mattered — is gone.

Our own runs store the prompt list verbatim in every record, and the comparison tool refuses to present two runs as like-for-like if the lists differ. That is the behaviour to demand from any tool you buy.

The flaw we found in our own set

We track eight questions and we know what is wrong with them, so here it is. In the Philippines, where we are based, “AEO” usually means Authorized Economic Operator — a customs-accreditation programme run by the Bureau of Customs. Ask a local-sounding AEO question and the sources returned are about customs brokerage, not marketing.

That acronym clash is documented and real, and none of our eight questions tests it. All eight are English-phrased and market-neutral. So our published score says nothing at all about the Philippine market, and for a while several of our own pages implied it did. We corrected them and we are adding a Philippine question to the next round.

The general lesson is worth more than our specific mistake: write down what your set cannot see. Ours also excludes every non-English phrasing, every vertical-specific question, and every question a buyer asks after a demo. Those are deliberate omissions, and naming them stops a limited number being read as a complete one.

Our eight, as an example

Published in full on our score page, with every answer stored. They cover the category question, the tools question, two price and budget questions, a how-to, a citation-tracking question and a competitor-alternatives question. Five of the five types, unevenly. You are welcome to steal the structure.

Common questions

Should I track my brand name in AI visibility tools?

Not as part of a visibility score. A prompt containing your brand name measures whether an assistant can look you up, which it almost always can, for a question almost nobody asks — and it will flatter the report while telling you nothing. Checking that an assistant describes you accurately is a worthwhile but separate job, and it should be reported separately.

How many questions should I track?

Eight to fifteen for a single product in a single market: enough for a percentage to mean something and few enough that you can read every answer by hand, which you should for the first few months. Around 25 once you have several segments or a second market. A hundred or more only if a team is acting on the results, because a large set nobody reads is just a bigger bill.

Can I change my tracked questions later?

You can add, but you should never swap or reword. Every metric is computed over the question set, so changing it changes the denominator and month two stops being comparable with month one. Version the list with a date, store it with every run, and treat new questions as a separate series until they have their own history.

What is wrong with Bungad's own eight questions?

They are all English-phrased and market-neutral, so they test nothing about the Philippine market where we are based — and in the Philippines AEO usually means Authorized Economic Operator, a customs programme, which returns an entirely different set of sources. That acronym clash is real and documented, none of our eight questions tests it, and for a while several of our own pages implied otherwise. We corrected them and are adding a Philippine question to the next round.

Try three of them

Enter your domain and we will put three buyer questions to Perplexity and show you who gets named and cited. A reasonable way to sanity-check a list before you commit to it.

3 questions on Perplexity, free. A domain checked in the last week is served from that stored run rather than asked again, and the daily free allowance resets at midnight UTC.