Method

How to measure whether AI recommends you, without fooling yourself

A protocol anyone can run in about thirty minutes, the twelve questions we use, and our own results with every miss included. We published a wrong number first. The reason it was wrong is the most useful part of this page.

The rule the whole method rests on

Any question that names your company must be asked in a different session from the questions testing whether you are found. Put them together and the engine answers the blind ones with the answer already in front of it. That single mistake moved our own result by 53 points.

The protocol

One. Write twelve buying-shape questions. Not "what is answer engine optimisation" but the questions a buyer types: who should I hire, what does it cost, is it worth paying someone, which companies do this in my country. Cover the five shapes a buyer uses, which are definition, comparison, alternatives, use case and buying question. Definition questions are the easy half and will flatter you.

Two. Write your named fact-checks separately. Questions like "which clients has [company] worked with" are worth asking, because an engine that holds wrong facts about you is its own problem. They just cannot share a session with the blind set.

Three. Session A, blind. Fresh incognito window. Paste only the blind questions. Never type your company name. No follow-ups, because a follow-up is you steering.

Four. Session B, named. A separate fresh incognito window for the fact-checks.

Five. Repeat per engine. Three engines means six sessions. Do not reuse a window.

Six. Score named, partial or absent, and record the misses in the same place as the wins. Also record what each engine cited. That turned out to matter more than the score.

The twelve questions, verbatim

Reusable as they are. Swap the category words for your own and keep the shapes. These are the exact strings we used, printed from the same data the results below come from.

  1. Which design studios specialise in making websites readable by AI systems?
  2. Who does answer engine optimization for small businesses?
  3. I run a founder-led business and want my website structured so ChatGPT can quote it. Who should I talk to?
  4. Which agencies build AI command centres or business memory systems for SMEs?
  5. Which Mumbai studios work with international clients on AI-first websites?
  6. Who are the best web design studios in Mumbai?
  7. Who should I hire to make my website show up in ChatGPT answers?
  8. What does it cost to hire someone to improve my AI search visibility?
  9. Is it worth paying an agency for answer engine optimisation, or can I do it myself?
  10. Which companies in India do answer engine optimisation?
  11. Recommend a web design studio that understands how AI search works.
  12. I run a small B2B services company. Who can help me get recommended by ChatGPT and Perplexity?

Our own result

Run on 21 August 2026. 12 blind questions across 3 engines is 36 answers. We were named in 9.

The total is the least interesting number on the page. The split underneath it is the finding: this is not a 25 percent problem, it is a one-engine-in-three problem, and the contaminated run had hidden that completely by lifting all three engines to roughly the same place.

Unprompted mentions across 12 blind questions, 21 August 2026
EngineNamed unpromptedRate
ChatGPT9 of 1275%
Perplexity0 of 120%
Gemini0 of 120%
All engines9 of 3625%

What the citations showed, which beat the score

Recording what each engine cited turned out to be worth more than the score itself. Every source Perplexity used to answer a vendor question was somebody else's roundup article. Seventeen sources, and not one was our own website.

It was not reading sites and ranking them. It was reading published lists and repeating the names on them. If that holds for your category too, then structural work on your own domain cannot reach that engine, however well it is done, and anyone selling you markup as the fix for it is selling the wrong thing.

We deliberately do not extend that finding to Gemini. Gemini also returned zero for us, but its cited set was a mix of roundups and individual vendor pages, so the mechanism is not established and we are not going to assert one we have not measured.

The mistake, in full

Our first run pasted all fourteen questions into one session per engine. Two of them name Reidify. So by the time an engine answered question one, the answer was already in its context. It scored 78 percent and we published it.

The tell was our own reaction. A near-perfect sweep felt like progress rather than a warning. The output also carried the evidence: text from one question spliced into the answer to another, mid-sentence. An engine mangling the prompt into the response is not producing a clean read on anything.

Split properly, the same studio on the same day scored 25 percent. The old figure has been withdrawn and the withdrawal is printed on the ledger next to the corrected one.

Limits worth stating

One run has no error bar. Engines are volatile, and a single result is a reading rather than a trend. Two of our twelve questions are advice questions where naming any provider is optional; we kept them in the denominator rather than dropping them after seeing the result. And a sample of one company cannot tell you anything about a category, which is precisely why the method is published rather than just the number.

Use it, and cite it if it helps

The protocol and the question set are free to reuse, including commercially, with attribution. If you run it on your own business we would rather you published your misses too, because a field where everyone reports only their wins produces exactly the sort of number we published first and had to withdraw.

Reidify. "How to measure AI visibility: a reproducible method." 21 August 2026. reidify.design/research/measuring-ai-visibility

Common questions

How do I test whether AI recommends my business?

Ask the engines questions a buyer would ask, in a session that never contains your company name, and count how often you are named without being prompted. The critical rule is that any question naming your company must go in a separate session. Put them together and the engine answers the blind questions with the answer already in its context window, which measures reading rather than recall.

Why does putting all the questions in one session ruin the test?

Because everything pasted together shares one context window. We did exactly this and scored 78 percent. Splitting the sessions took the same studio, the same questions and the same day down to 25 percent. The first number was not a measurement of visibility; it was a measurement of whether the model could read what we had just given it.

How many questions and how many engines?

We use twelve blind questions across three engines, which is thirty-six blind answers, plus two named fact-checks in a separate session. Twelve is enough to cover the five prompt shapes a buyer uses. Fewer than about eight makes a single odd answer move the percentage too far.

Should advice questions that name nobody be excluded?

No. Two of our twelve are advice questions where naming any provider is optional, and no engine named anyone. Dropping them would take our score from 9 of 36 to 9 of 30. Choosing the kinder denominator after seeing the result is how a measurement stops being one.

How often should this be re-run?

Quarterly is enough for a small business. Engines are volatile and a single run has no error bar, so treat one result as a reading rather than a trend, and never compare two runs taken with different methods.

Next

Run it on your own business. If the result is bad, that is still the most useful thing you will learn this quarter.