On this page
- 01The short answer
- 02Define what visibility should make possible
- 03Start with evidence the platforms expose directly
- 04Use a stable prompt portfolio to observe what reports cannot show
- 05Keep mentions, citations, coverage and context distinct
- 06Connect citations to site behaviour and qualified outcomes
- 07Compare like with like and annotate the changes
- 08Build a report that separates fact, interpretation and action
- 09Use measurement to run focused improvements
- 10Frequently asked questions
Measure AI search visibility through four separate evidence layers: official platform data, repeatable prompt observations, referral behaviour and business outcomes. Report mentions, citations, source patterns and answer context independently. There is no reliable universal position that can combine Google AI features, ChatGPT, Copilot, Gemini and Perplexity into one rank.
Traditional rank tracking usually associates a URL with a query and position. Generated answers can synthesize several sources, mention a brand without linking to it, cite a page without recommending the company or change between observations. Measurement therefore needs more than a renamed ranking chart.
Begin with an AI search visibility audit to define audiences, decisions and the baseline. This guide explains how to maintain the measurement system afterwards and how to connect it to Qreativa’s AI SEO service.
Platform evidence
Use impressions, cited pages and grounding queries where official reports provide them.
Prompt observations
Track presence, citations and framing across a fixed, representative prompt set.
Audience response
Identify referrals, useful site journeys and demand that can be observed.
Business impact
Connect qualified actions and revenue carefully, preserving attribution limits.
Define what visibility should make possible
A useful scorecard begins with the business decision. A publisher may want its expert analysis to be discovered more often; a B2B company may want to enter vendor shortlists; an ecommerce brand may need accurate product comparisons. The same mention can have different value in each case. Write the priority audience, decision and expected next step before selecting metrics.
Discovery
Can the brand or its expertise become visible while the audience learns about a problem?
Consideration
Does the brand appear in relevant shortlists or comparisons with accurate context?
Verification
Can people follow a citation and confirm the claim through a credible source?
Action
Does visibility contribute to a useful visit, enquiry, purchase or assisted decision?
Start with evidence the platforms expose directly
Google Search Console provides a dedicated view for impressions and pages shown in generative features in Search and Discover, with country, device and time dimensions where supported. Bing Webmaster Tools reports total citations, average cited pages, grounding queries, page-level citation activity and trends across Microsoft AI experiences. Keep each platform’s definitions alongside the data.
| Evidence | What it supports | What it does not prove |
|---|---|---|
| Google generative-feature impressions | How often site URLs appeared in the reported Google AI experiences | A fixed rank, a citation in every answer or the reason a page was selected |
| Bing total citations | How often site content was displayed as a source in supported AI answers | Placement, authority, sentiment or importance inside an individual answer |
| Bing grounding queries | Sampled phrases used to retrieve cited content | A complete prompt list or direct record of every user question |
| Cited pages | Which owned URLs are being used as sources | Whether the brand itself was recommended or converted a user |
Annotate product changes, report availability and methodology changes. A new reporting interface can create an apparent trend even when customer behaviour has not changed. Do not combine metrics from different platforms simply because each column contains a number.
Use a stable prompt portfolio to observe what reports cannot show
Controlled prompt monitoring can show whether the brand is mentioned, how it is described, which competitors appear and which sources are visible. Use a representative prompt portfolio grouped by audience and decision stage. Freeze a core set for trend comparison and maintain a smaller exploratory set for new questions. Do not silently replace underperforming prompts.
Save the observation conditions
- Exact prompt wording and prompt-group identifier
- Platform, product surface and model label where visible
- Country, language, device and account state
- Date, time and collection method
- Brand and competitor mentions with surrounding context
- Every visible cited URL and its source type
- Answer limitations, refusals or absent citations
If prompts are repeated, separate repeated samples from distinct customer questions. A hundred runs of the same prompt do not represent a hundred people. Use repetition to assess variability, not to manufacture a larger market sample.
Keep mentions, citations, coverage and context distinct
| Metric | Working definition | Reporting caution |
|---|---|---|
| Prompt coverage | Share of priority prompts where the brand has a relevant presence | State the prompt set; changing it changes the denominator |
| Mention rate | Share of observed answers that name the brand in any context | A mention may be irrelevant, negative or factually wrong |
| Citation rate | Share of observed answers that link or attribute a source connected to the brand | Define whether owned and third-party sources are reported separately |
| Source distribution | Which owned and independent domains recur across citations | Frequency is not a platform-defined authority score |
| Answer context | How accurately and usefully the brand is described or compared | Requires a documented rubric and human review |
| Competitive presence | How often named alternatives appear for the same prompt set | Do not infer causation from co-occurrence alone |
For qualitative context, use a small rubric with explicit categories: accurate, partly accurate, materially misleading or not assessable. Record the evidence behind the classification. Automated sentiment alone can miss an answer that sounds positive but assigns the brand to the wrong category.
Connect citations to site behaviour and qualified outcomes
OpenAI states that referral URLs from ChatGPT search include utm_source=chatgpt.com, which can be analysed in web analytics. Also review referrer data and server logs where lawful and appropriate, because attribution can be reduced by consent choices, browser behaviour and cross-device journeys. Preserve the difference between a cited page, a visit and a commercial action.
Citation
The platform displayed a page as a source; no visit is implied.
Referral
A user reached the site from an identifiable AI-search source.
Qualified action
The visit produced a relevant enquiry, signup, purchase or useful next step.
Business contribution
The interaction contributed to an outcome, with attribution limits stated.
Compare landing-page engagement, assisted journeys, lead quality and sales context rather than celebrating traffic in isolation. For low-volume B2B journeys, qualitative sales evidence may be useful, but it should supplement recorded analytics rather than overwrite them.
Compare like with like and annotate the changes
Preserve the original prompt set, source definitions and platform conditions so later results remain comparable. Report both the stable core and any newly added exploratory prompts. Segment by market, language, topic and decision stage before interpreting a total. A global improvement can hide a decline in the market that actually matters.
Annotate material changes
- Platform interface, model or reporting changes
- Website launches, migrations and crawler-policy updates
- New or substantially revised source pages
- Digital PR coverage and independent reviews
- Product, offer, price and availability changes
- Seasonality, news events and competitor activity
- Changes to analytics, consent or conversion definitions
Use ranges and confidence labels when the sample is small or volatile. A five-point movement based on four observed answers is not equivalent to a five-point movement in a stable platform report covering thousands of impressions.
Build a report that separates fact, interpretation and action
| Report layer | Include | Decision it supports |
|---|---|---|
| Platform evidence | Impressions, cited URLs, citations, grounding-query samples and trends | Where owned content is observably surfaced |
| Prompt observations | Coverage, mentions, citations, context and competitors for the fixed set | Which decision areas require investigation |
| Source analysis | Owned and third-party domains, evidence type, accuracy and freshness | Which sources should be improved, earned or corrected |
| Audience response | Referrals, landing-page behaviour and qualified actions | Whether visible answers lead to a useful journey |
| Business context | Lead quality, assisted revenue and sales observations | Whether further investment is commercially justified |
For every conclusion, show the underlying source and label whether it is directly observed, supported by platform documentation or inferred. Then attach an owner and next review date. Reporting should make uncertainty easier to see, not bury it beneath a polished dashboard.
Use measurement to run focused improvements
Select one evidence-backed constraint at a time: blocked access, an inaccurate company description, a weak comparison page, missing first-party proof or an important topic with no credible source. Record the intervention, expected signal and evaluation window before implementation. Avoid rewriting every page for AI at once; that destroys the baseline and makes learning harder.
Weekly
Watch access failures, material brand errors and reporting anomalies that need prompt action.
Monthly
Review platform trends, core prompt observations, cited pages and qualified referrals.
Quarterly
Reassess prompt coverage, competitors, source gaps, business contribution and priorities.
After major change
Repeat the relevant baseline after launches, migrations, rebrands or substantial product updates.
The cadence should reflect commercial importance and available evidence. Stable low-volume topics do not need daily prompting; a critical inaccurate answer about a regulated or high-risk product may require immediate investigation.
AI search measurement FAQs
Is there a single AI search ranking position?
No reliable universal position spans all generative search products. Each platform exposes different data, and a generated answer may mention a brand, cite several pages or change between observations. Report the platform, prompt set, market, date and metric definition rather than presenting one composite rank.
What is the difference between mention rate and citation rate?
Mention rate records how often the brand is named within the observed answer set. Citation rate records how often an answer links or attributes a source connected to the brand under the stated definition. A brand can be mentioned without a citation or have an owned page cited without being recommended.
Can ChatGPT referral traffic be tracked in analytics?
OpenAI says ChatGPT search referral URLs include utm_source=chatgpt.com. Analytics can therefore identify some visits, although consent, browsers, applications and cross-device behaviour can create gaps. Referral traffic measures visits, not every answer in which a source appeared.
How should competitors be included in AI visibility reporting?
Use the same fixed prompt portfolio and observation rules for every named competitor. Report presence and supporting sources by decision group rather than using one total to imply market share. Competitor comparison is diagnostic; it does not reveal a platform’s private ranking logic.
How long should an AI search measurement baseline run?
The period should be long enough to capture the platform data and repeated observations required by the decision. A monthly baseline is often more defensible than a single-day snapshot, but fast-moving launches or material errors may justify shorter diagnostic checks. State the period and sample size in every report.
Can AI search visibility be tied directly to revenue?
Sometimes a referral and subsequent outcome can be observed, but many journeys are assisted, cross-device or completed without a click. Connect qualified visits, conversions, CRM outcomes and sales evidence while preserving attribution limits. Do not assign all later revenue to an earlier mention.
Michele Eccher


