Short answer

A real GEO audit tests your brand across all seven major AI answer surfaces — ChatGPT, Claude, Gemini, Perplexity, Grok, Microsoft Copilot, and Google AI Overviews — using repeatable prompt sets per surface, not a single spot-check on one app. Each engine needs a different testing approach (web UI vs. API, logged-in vs. logged-out, single-shot vs. repeated runs), and each industry needs different prompt categories: fintech audits add a compliance-citation check, France DPE audits add French-language and local-registry checks, and local service audits add location-anchored "near me" prompts. Run it quarterly, log every answer verbatim, and track whether your brand is named, how it's described, and which competitors get cited instead.

Why a single-engine check isn't a GEO audit

It's tempting to open ChatGPT, type your brand name, feel good or bad about the answer, and call it a day. That's a screenshot, not an audit. The problem is that AI answer engines disagree with each other constantly — a brand cited confidently in Perplexity can be missing entirely from Gemini, and Google AI Overviews frequently pulls from a different source pool than Copilot even though both sit on Microsoft's Bing index underneath. If you only check one surface, you're extrapolating from a sample size of one.

The seven engines below account for the overwhelming majority of AI-driven research and purchase-influence traffic in 2026: ChatGPT alone passed a billion monthly active users, Google's AI Overviews now surface on a majority of informational queries, and Perplexity, Claude, Gemini, Grok, and Copilot each carry meaningfully different user bases and citation behavior. A GEO audit that skips any of them is guessing about the gaps.

There's also a structural reason multi-engine testing matters more than it used to: these models retrieve differently. Some browse live search results at answer time (Perplexity, Copilot, AI Overviews, and browsing-enabled ChatGPT), some lean more heavily on what was baked into training plus a lighter retrieval layer (base Claude and Gemini responses without an explicit search step), and Grok blends its own model with real-time X data. Testing only one engine tells you about one retrieval pattern — not about how discoverable you are as a source in general.

The audit method, engine by engine

Below is exactly how to test each of the seven engines: what interface to use, how many times to run each prompt, and a sample prompt to start from. Run every prompt logged out (to see the default, unpersonalized answer) and, where relevant, logged in with a plan that has browsing enabled — the two can differ substantially.

Engine 1 of 7

ChatGPT

Test in the consumer web app (chat.openai.com) with browsing/search mode on, since that's what most buyers actually use — a raw API call without the "web" tool often skips live retrieval and answers from training data alone, which understates or overstates your visibility depending on how recent your content is. Run each prompt fresh (new chat, no memory) three times to catch variance, then once more with ChatGPT's memory/personalization on if your team already uses it daily, since personalization can shift which sources get surfaced.

"What is [Brand] and what do they do?" · "Best [category] for [use case]?" · "[Brand] vs [Competitor] — which is better for [use case]?"

Log whether your brand is named unprompted in the category question, what it says about you in the direct question, and whether the comparison question cites your own site or a third-party review instead.

Engine 2 of 7

Claude

Test in claude.ai with web search enabled (Claude added live browsing as a toggled feature — confirm it's on, since the default answer without it draws only on training knowledge and general knowledge cutoff, which will make your recent content invisible even if it's well optimized). Claude tends to be more conservative about naming brands it can't verify, so a "no mention" result here is a stronger signal of a genuine visibility gap than the same result on a more citation-eager engine.

"I'm evaluating [category] providers — walk me through what [Brand] offers and how it compares to alternatives." · "What do independent sources say about [Brand]?"

Because Claude explains its reasoning more often than other engines, read the full answer for *why* it did or didn't cite you — it frequently states directly that it lacks recent information, which tells you exactly what content gap to fill.

Engine 3 of 7

Gemini

Test both the standalone Gemini app and Gemini inside Google Search (the "AI Mode" tab), since Google routes the same query differently depending on entry point and the two can produce different citation sets. Gemini draws heavily on Google's own index and Knowledge Graph, so a strong Google Business Profile, Wikipedia presence, and structured data footprint move the needle here more than on other engines.

"[Category] options in [location/niche] — compare the top choices." · "Is [Brand] a good fit for [use case]?"

Check whether Gemini's answer matches what your Google Business Profile and Knowledge Panel (if you have one) say — mismatches are usually the fastest fix available.

Engine 4 of 7

Perplexity

Perplexity is search-first by design, so every answer already browses live and shows its numbered sources inline — this makes it the easiest engine to audit precisely, since you can click through to see exactly which pages it pulled from. Test in the web app with "Auto" or "Best" model selection (matches what most free users see) and repeat with Pro search enabled if your buyers are likely to be paid users. The API mirrors app behavior closely enough to script larger prompt batches here if you need scale.

"[Category] for [audience] — what are the top options and why?" · "Reviews and reputation of [Brand]"

Record the exact source URLs cited, not just whether you were mentioned — if a competitor's blog post is cited instead of your own page on the same topic, that's your specific content gap to close.

Engine 5 of 7

Grok

Test inside X (formerly Twitter) where Grok is embedded, and separately at grok.com if your buyers might reach it there. Grok has real-time access to X posts and trending discussion, so it can surface reputation signals — complaints, praise, recent news — that other engines miss entirely because they don't index social conversation as heavily. This makes Grok audits partly a social-listening exercise: what people are actually saying about you on X shows up here fast.

"What's the current sentiment around [Brand] on X?" · "Compare [Brand] to [Competitor] for [use case] — what are people saying?"

If Grok surfaces old complaints or outdated info more than your competitors' audits show, that's a signal your X presence and public reputation management need attention, not just your content.

Engine 6 of 7

Microsoft Copilot

Test in the Copilot web app (copilot.microsoft.com) and, if your buyers are enterprise users, inside Microsoft 365 Copilot where the answer can blend web results with a company's internal tenant data — audit the public web-facing version unless you're specifically checking how your own org's content surfaces internally. Copilot runs on Bing's index, so a healthy Bing Webmaster Tools presence and Bing-indexed schema markup matter more here than for Google-anchored engines.

"Summarize what [Brand] does and how it's positioned in [category]." · "What are the leading [category] providers and their strengths?"

Because Copilot and Bing share an index, a quick check of how your site performs in plain Bing search is a fast leading indicator of how Copilot is likely to treat you.

Engine 7 of 7

Google AI Overviews

There's no way to script this one — test by hand, logged out, in a standard Google Search (not the separate AI Mode tab, which is really Gemini; see above). AI Overviews only appear for a subset of queries, so your first job is simply determining whether your target prompts trigger an Overview at all — many transactional and branded queries still don't. When an Overview does appear, note the exact cited sources shown in the expandable source list, since that list is Google's closest thing to a public citation audit trail.

"what is [category] and how does it work" · "[category] near me" (for location-based businesses) · "how to choose a [category] provider"

Because Overviews are query-triggered rather than always-on, test question-shaped, informational phrasing — not just brand-name lookups — since that's what actually earns an Overview appearance.

Engine comparison table: how to actually test each one

Access method, limits, and testing notes for all seven engines
Engine How to test Public API? Practical limits What to prompt
ChatGPT Web app, browsing mode on Yes — but API calls skip live browsing unless the web tool is explicitly enabled Usage caps on free tier; answers vary run to run Direct brand, category "best of," head-to-head vs. named competitor
Claude claude.ai, web search toggle on Yes — same caveat, browsing is a tool call, not default Conservative citation behavior can undercount true visibility Evaluation-style prompts, "what do independent sources say"
Gemini Gemini app and Google AI Mode tab (test both) Yes, via Google AI Studio / Vertex App vs. Search-embedded answers can diverge Location/niche comparisons, "is X a good fit"
Perplexity Web app, Auto and Pro search modes Yes — API behavior tracks the app closely Rate limits on API scale for large prompt batches Category roundups, reputation/review queries
Grok Inside X and at grok.com Yes, via X's developer platform Answers weighted toward real-time X discussion Sentiment checks, "what are people saying"
Microsoft Copilot Web app; M365 Copilot separately for enterprise No public generation API for web Copilot Runs on Bing's index — test manually, logged in and out Positioning summaries, category-leader questions
Google AI Overviews Standard Google Search, logged out No public API Only triggers on a subset of queries — manual testing only Informational, question-shaped, and "near me" queries

Industry playbooks: what changes by vertical

The seven-engine method above is the constant. What changes by industry is the prompt set and the review layer you apply to the answers. Three verticals need materially different audits from a standard B2B company: fintech, France's DPE (diagnostic de performance énergétique) providers, and local service businesses. Here's what's different for each.

Fintech: compliance-safe prompts and regulatory citation checks

A fintech GEO audit needs a compliance pass that a typical B2B audit doesn't. Run your standard prompt set, then add a second layer that checks two things: whether AI engines cite regulator-approved or regulator-adjacent sources (SEC filings, FINRA BrokerCheck, CFPB complaint data, or your jurisdiction's equivalent) rather than a marketing page when describing your product, and whether any AI-generated summary of your offering implies investment, credit, or financial advice that hasn't been reviewed by legal or compliance. An AI engine paraphrasing your APY or fee structure incorrectly is a real-money problem, not just a visibility gap.

  • Prompt for regulatory standing directly: "Is [Brand] a licensed/regulated [product type] provider?" and check what the engine cites as proof.
  • Test disclosure-sensitive prompts: "What are the fees/rates for [Brand]'s [product]?" — flag any answer with numbers you haven't published or approved.
  • Check whether engines default to citing third-party comparison sites (which may carry stale or inaccurate rate data) over your own current disclosures.

France DPE providers: French-language testing and local regulatory sources

Every prompt in the core method needs to run twice for a French DPE (diagnostic de performance énergétique) provider: once in English and once in French, since AI engines frequently answer in the query's language and draw from a different citation pool per language — an engine might cite a solid French source in a French-language answer while defaulting to a generic English explainer when asked in English. The second layer to check is regulatory currency: DPE methodology changed materially in 2021, and older or lower-quality sources still circulate outdated calculation methods, so audit whether engines cite ADEME (the official French agency) and reference current diagnostiqueur certification rules rather than pre-reform guidance.

  • Run prompts in French: "Qu'est-ce qu'un DPE et comment [Brand] peut-il m'aider ?" alongside the English equivalent, and compare which sources each language surfaces.
  • Test department- and city-level queries — "diagnostiqueur DPE à [ville]" — since DPE is delivered hyper-locally and generic national prompts miss most real buyer phrasing.
  • Check citation currency: does the answer reference the post-2021 DPE methodology, and does it name ADEME or an equivalent official source rather than an unverified blog?

Local service businesses: location-based and "near me" prompt testing

A local service business audit lives or dies on location-anchored prompts, because that's how real buyers actually search — almost nobody searches a plumber, electrician, or HVAC company by brand name first. Test "near me" phrasing, specific neighborhoods or cities, and service-area language, and pay close attention to Google AI Overviews and Gemini specifically, since both draw heavily on Google Business Profile data, local reviews, and map-pack signals that barely factor into ChatGPT or Claude's answers. A local audit that only runs generic category prompts will systematically miss the queries that drive real leads.

  • Run true "near me" and city-anchored prompts: "[service] near me," "best [service] in [city/neighborhood]," "[service] that serves [zip code/area]."
  • Check whether the engine surfaces your Google Business Profile details (hours, service area, review count) accurately — mismatches here often come from an out-of-date profile, not a content problem.
  • Test emergency/urgency phrasing common to local trades: "24-hour [service]," "same-day [service] near me" — these prompts convert differently and are worth auditing separately from general category questions.

Running the audit end to end

Put it together as a single repeatable process: build your prompt matrix (core brand, category, and comparison prompts, plus any industry-specific additions above), run every prompt on every relevant engine at least three times to average out response variance, and log the results in a spreadsheet with columns for engine, prompt, whether the brand was named, what was said, and which sources were cited. Score each cell as named-with-accurate-info, named-with-inaccurate-info, or not named — that three-way split is more actionable than a single visibility percentage, because "named but wrong" and "not named" call for completely different fixes.

From there, prioritize by where the gap is largest and the buyer intent is highest — a missing citation on a bottom-of-funnel comparison prompt matters more than one on a broad awareness question. For the baseline checklist of what content, schema, and entity signals to check before you even get to engine-by-engine testing, see GEO Audit for B2B Companies: What to Check First — that page covers the foundational audit; this one covers the full multi-engine, multi-industry testing method that sits on top of it.

The Arrow read

Seven engines is the floor, not the ceiling.

Manually running this audit across seven engines, three industries, and dozens of prompt variants is a real time investment — and it goes stale every time a model updates. That's exactly what Arrow's GEO platform automates: continuous prompt testing across all seven surfaces, tracked over time, with the same three-way named/inaccurate/absent scoring described above. For the general audit foundation, read GEO Audit for B2B Companies: What to Check First, and for what a genuinely serious baseline audit includes end to end, see The AI Visibility Audit: What a Serious Baseline Includes. Or skip the manual spreadsheet entirely and run the free GEO audit to see where you stand today.