How we measure AI visibility.

Method updated September 9, 2026 · the weekly pilot

We track answers to an agreed set of buying questions, preserve the evidence, and turn what we find into one page task each week. This page explains what those observations can tell you and where their limits are.

See a worked weekly brief · View the $990 pilot

1. Agree the questions before measuring.

At kickoff, which happens before your first payment, we confirm your buyer segment and the country or region its buyers are in, product names and aliases, five competitors, and 50 questions. The starting mix is 30 discovery and alternatives questions plus 20 branded facts and comparisons. We agree the wording with your team before collecting the baseline, and the agreed set goes into your order with the date of each delivery.

Questions that name your brand are kept separate from questions where buyers are choosing a vendor. We do not replace questions simply because their answers look better. Changes to the question set or measurement settings require a new baseline.

2. Repeat the API observations and keep the answers.

The weekly protocol repeats each question three times on each of the OpenAI, Google Gemini and Perplexity APIs: 150 planned answers per engine, 450 in total. Search is enabled in the API requests; the archive records the citations actually returned. Search availability does not guarantee that every answer searches or cites a page.

Gemini runs with Google Search grounding switched on, so those answers are Google’s AI answering out of Google’s own index, and the archive records the citations it returned. OpenAI and Perplexity run with their own web search enabled on the same questions, in the same week, so the three are directly comparable.

Three repeats help expose variation; they do not establish statistical significance. Model and search behaviour can also change without a visible change in the model name.

3. Keep different signals separate.

These are observations from a fixed question set. They are not market share, search volume or a count of actual buyers.

4. Compare only compatible runs.

We compare against an earlier date with the same questions, buyer region, settings and reported models. At least 95% of planned discovery/alternatives observations must have successful matching answers in both runs. Otherwise, the report withholds the change and explains why.

Changes use the matched observations and are expressed in percentage points. The 95% rule is a coverage check, not a 95% confidence level or a significance test. Even a compatible before-and-after change does not prove that a page edit caused it.

5. Label actual ChatGPT and Gemini app checks separately.

Our recurring measurement uses APIs. A separate consumer-app check captures the first answer to a preselected question in a fresh conversation, with its date, exact prompt, displayed model or mode, search setting and citations where available. If the app does not display a detail, we record it as not shown.

App answers can differ with conversation history, personalisation, account settings, location and search behaviour. A US buying question entered from Korea is not a measurement of a US-located user's session. One missing mention does not establish that your brand is invisible to buyers.

We retain the full answer alongside representative screenshots and record appearances as well as absences. We do not repeat a query until it produces a more useful sales claim. App checks are not added to the API denominator.

6. Connect the evidence to a published change.

A missing AI mention alone does not prove a defect on your website. We check the relevant public page before proposing an edit. Each weekly task identifies the exact URL, evidence, edit or checklist, named owner, due date and remeasurement date. Your team approves and publishes the change.

We choose that one page the same way every week, and the brief says which rule picked it. Of the week's answers that left you out, we group them by the page a buyer would land on, and start with the page the largest group points to. We then read that page against the questions behind the group: if it already states the facts those questions asked for, it is not a candidate however often it came up, and we move to the next group. So a page is picked because many answers point at it and it is missing a fact that was asked for, not because it is weak in general. The edit we propose adds facts your team has already confirmed; where we cannot source a fact, we ask for it instead of writing one.

We record what shipped, check the published page, and review comparable observations. At the end of the first month, we discuss the remaining work and whether another paid month is useful. Ranking, traffic and revenue increases are not guaranteed.

About the historical research on our homepage.

The September 2026 exploratory sample used a different protocol from the current pilot. Its archived answers do not retain all model and request settings, and questions were repeated across vendors and engines. It is not a current customer baseline or a before-and-after service result.

The homepage task preview and worked weekly example are labelled examples, not customer results.

Email your buyer questions