Back to Blog

AI Search Visibility Audit: Measure Google, Bing, and ChatGPT

·13 min read·Rendframe·AI Search, SEO, Measurement, Content Strategy

AI-search visibility used to be measured with screenshots and anecdotes: ask a chatbot once, celebrate a mention, then sell a “GEO score.” That is no longer good enough. Google now reports generative-AI impressions in Search Console, Bing reports citations and grounding queries, and ChatGPT marks referral links. A business can finally build a useful baseline—provided it does not pretend these different signals are one ranking.

Editorial measurement loop connecting crawl access, useful evidence, AI citations, qualified visits, and business outcomes
AI visibility is an evidence chain: eligibility, useful source material, citations, qualified visits, and outcomes.

The practical approach is to audit three layers separately: can the systems access and index the right page, is the page useful enough to support an answer, and does that visibility create qualified demand? Fix the first broken layer. Do not buy content volume, special markup, or automated prompt tracking before you know which problem exists.

What changed in 2026

On 31 August 2026, Google completed the worldwide rollout of Search Console’s Generative AI performance report. It shows impressions for links to your site in AI Overviews and AI Mode, broken down by page, country, device, and date. It does not show a universal “AI rank,” and low-volume properties may not receive enough data to display the report.

Bing Webmaster Tools’ AI Performance takes a different view: total citations, average cited pages, sampled grounding-query phrases, URL-level citation activity, and trends across Microsoft Copilot, Bing AI summaries, and selected partner integrations. Microsoft explicitly warns that a citation count is not position, authority, or importance.

OpenAI’s current publisher guidance says any public site can appear in ChatGPT search, that OAI-SearchBot must not be blocked if content is to appear in summaries and snippets, and that outbound referrals include utm_source=chatgpt.com. That measures visits, not unclicked mentions.

These releases make AI discovery more observable, not deterministic. McKinsey’s July 2026 consumer research also found that generative AI is used for product research while remaining among the less-trusted sources it measured. Visibility without specific, verifiable evidence may win an impression and still lose the decision.

Preserve a baseline before changing the site

Choose one commercially important topic cluster: a service, problem, category, or location that can lead to a real enquiry or sale. Export 90 days of ordinary Search Console and analytics data. Export the new Google generative-AI report and Bing AI Performance where available. Save the exact date range and property settings.

Then make an inventory of five to fifteen pages that should answer the cluster. For each page record canonical URL, index state, title, last meaningful update, owner, conversion action, organic impressions, AI impressions or citations, ChatGPT referrals, and qualified outcomes. Zero is a valid baseline. “No report” is not zero: it can mean insufficient data or eligibility.

Separate four observations in every report:

SignalWhat it supportsWhat it cannot prove
AI impressionA link was displayed on a reported surfaceThat a person noticed, trusted, or clicked it
CitationA page was used as a displayed sourceIts exact influence or rank in the answer
Referral visitA click reached your site with source evidenceAll prior AI influence or no-click exposure
Lead or orderA business outcome occurredIncremental causation without a comparison

Layer 1: prove that the page is eligible and reachable

Start with the canonical page, not the homepage. Confirm it returns a successful response without login, challenge, or geography-dependent error; is not blocked by robots.txt; carries no unintended noindex or nosnippet; has a self-consistent canonical; appears in the XML sitemap; and is linked from a relevant page a visitor can find.

Google says a page must be indexed and eligible to show a snippet to appear as a supporting link in its generative features. For ChatGPT summaries, OpenAI names OAI-SearchBot. A correct robots rule is still insufficient if a CDN or WAF returns 403, a JavaScript shell contains no useful initial content, or the canonical points elsewhere. Check the response and rendered page with the real crawler user agent where lawful, then inspect server or edge logs.

Do not confuse search access with model training. OpenAI documents GPTBot separately from OAI-SearchBot. Google uses Googlebot for Search and documents other controls for other AI uses. Decide access by purpose and content zone; the detailed workflow is in our AI crawler policy guide.

Layer 2: create material worth citing

A technically perfect page can remain invisible because it adds nothing. Google’s 2026 guidance gives “non-commodity content” priority: first-hand experience, a distinct and defensible view, and useful information that could not be reproduced by summarising the same ten competitors. This is a better editorial test than forcing every paragraph into tiny “AI-ready chunks.”

For the selected topic, add the missing decision evidence:

  • a direct answer to the real question, followed by conditions and trade-offs;
  • specific scope: audience, market, date, product version, price basis, or eligibility;
  • original evidence such as a tested workflow, comparison method, failure analysis, calculator, annotated example, or disclosed dataset;
  • links to primary sources beside material claims and a visible review date;
  • a responsible author or organisation, with a clear way to verify who stands behind the page;
  • consistent facts across the page, structured data, profiles, product feeds, policies, and support content;
  • a next action that fits the question: check availability, calculate cost, inspect an example, book a review, or buy.

Use structured data when it accurately describes a supported page type. It can clarify a product, organisation, article, or local business and support ordinary rich results. There is no special AI-citation schema, and markup that contradicts visible content is a quality defect, not an optimization.

Layer 3: measure each surface on its own terms

SurfaceUseImportant limit
Google Generative AI reportImpressions by page, country, device, and dateNo universal AI rank; report can be absent at low volume
Bing AI PerformanceCitations, cited URLs, sampled grounding queries, trendsAggregated preview data; citation count is not answer position
ChatGPT referralsSessions carrying utm_source=chatgpt.comNo-click mentions and copied brand searches are invisible
Analytics and CRMQualified visits, enquiries, orders, margin, outcome qualityAttribution remains partial and cross-device paths can break

Build a weekly page-level table rather than one vanity score. Show counts, not only percentages. Annotate launches, migrations, tracking changes, major publicity, and seasonality. Compare the selected cluster with a similar untouched cluster where practical. The business question is not “did citations rise?” but “did more relevant people reach a useful decision, and at what cost?”

Use manual prompt sampling as research, not telemetry

Platform reports do not reveal every question. Create 15–25 prompts from sales calls, support tickets, Search Console queries, onsite search, reviews, and customer interviews. Cover problem diagnosis, comparison, constraints, local availability, risk, price, and implementation. Translate intent for each market instead of translating a keyword list.

Record platform, model or mode, account state, country, language, date, exact prompt, answer, cited URLs, brand treatment, factual errors, and whether a useful next action was possible. Repeat a small sample on different days. Outputs vary with context and product changes, so never present one run as a stable market share.

Use the sample to find answer gaps and misinformation. Do not automatically generate a landing page for every prompt. One strong page can answer a family of related questions more coherently than twenty near-duplicates.

A 100-point audit that ends in a backlog

AreaPointsPass condition
Access and index eligibility20Priority pages are reachable, canonical, indexable, and snippet-eligible
Intent and information architecture15Each valuable question has a clear page and internal path
Original decision evidence25Pages contain specific, sourced, maintained material worth using
Entity and fact consistency15Names, offers, policies, locations, and markup agree
Platform measurement15Google, Bing, referral, and manual evidence stay separate
Business outcomes10Qualified enquiries or orders connect to page and source evidence

Score 0 for absent, half points for unreliable or partial, and full points for tested and owned. The thresholds are a planning heuristic, not a ranking model. Every lost point must become a specific fix with an owner, evidence of completion, and review date.

Four GEO shortcuts to decline

  • “Install llms.txt and rank in Google AI.” Google says Search ignores llms.txt. Maintain it only for a documented consumer, not as a magic signal.
  • “Add FAQ schema everywhere.” Unsupported or invisible markup does not create expertise. Use valid schema for its documented purpose.
  • “Publish every fan-out query.” Google warns that scaled pages made to manipulate results can violate spam policy. Cover the user’s task, not a synthetic keyword matrix.
  • “Our tracker says we rank third in ChatGPT.” A repeatable prompt panel can be useful research, but it is not a stable rank index or proof of causation.

A 30-day implementation plan

Days 1–5BaselineOne intent cluster, exports, priority URLs, outcomes
Days 6–10EligibilityIndex, snippets, bots, WAF, canonicals, internal links
Days 11–20EvidenceOriginal proof, primary sources, consistency, useful action
Days 21–30MeasureRe-export, prompt sample, leads, decision, next test

Do not expect indexing, citations, or meaningful demand to settle in 30 days. The sprint creates a trustworthy system and one controlled improvement. Continue only when the next experiment follows from observed evidence.

Frequently asked questions

How do I check whether my site appears in AI search?

Combine Google’s generative-AI impressions, Bing’s citation data, ChatGPT-tagged referrals, and a documented prompt sample. Each sees a different part of the journey.

Do I need llms.txt for Google AI Overviews?

No. Google says it does not use llms.txt for Search and requires no special AI file or schema. It may still be useful to another system that explicitly documents support.

Is structured data an AI ranking factor?

There is no special AI-search schema. Use accurate supported markup to describe the visible page and qualify for documented search features; do not promise citations.

Should we create content for every prompt?

No. Group related decision intent into a useful, original page. Update or create content only where customer evidence shows a real unanswered task.

What is the best AI-search KPI?

There is no single one. Track eligible pages, impressions or citations, qualified visits, and business outcomes as a funnel, with counts and limitations visible.

Sources and verification date

Verified 5 September 2026 against Google’s generative-AI optimization guide and Generative AI performance report documentation; Microsoft’s Bing AI Performance announcement; OpenAI’s publisher and developer FAQ; and McKinsey’s 2026 consumer trust analysis. Availability, interfaces, and reporting definitions can change; verify them in the linked products before acting.

Rendframe can turn this audit into a tested acquisition system: crawl diagnostics, information architecture, evidence-led content, analytics, and lead reconciliation. Review our AI attribution framework, see product engineering, or bring us one priority service and ten real customer questions for a focused visibility review.