Track AI visibility by preserving actual answers and source links for a defined question panel, then classifying mentions, citations and recommendations separately. Report the question set, surface, market, dates and counts. This page defines Arrow's proposed reference method; it does not certify that every product report or historical result was collected under this protocol.
Define the unit and the question cohorts
An observation is one attempted question on one specified surface in one session. Record the attempt even when it fails. A valid observation is a completed result under the defined collection conditions; technical failures remain excluded from outcome-rate denominators and are reported separately.
Create commercial questions about choosing a solution and informational questions about understanding a topic. Keep company-named identity checks in a third cohort. Recommendations prompted by your own brand name must not inflate the commercial discovery rate.
Preserve the wording, question ID and panel version. Adding a question changes the sample, so report additions separately from the original trend.
Keep collection environments distinct
Record the public interface or API configuration, search mode, language, known location settings, timestamp and session context. Use consistent conditions for repeated comparisons and record unavailable details as unknown. A prompt's requested country does not prove the actual search location.
An API answer is evidence about that configured API call, not automatically the corresponding consumer search product. Similarly, Google AI Mode and AI Overviews are separate surfaces. Keep their results separate even when the questions overlap.
Schedule repetitions across several days. Define retry rules before collecting. Retain the failed attempt and mark its replacement so that a retry does not become an extra success. Multiple completions of the same question are repeated observations, not independent buyers.
Apply a classification that separates outcomes
Check the company's identity using its domain and context. A shared name alone is insufficient. Preserve ambiguous cases for review rather than assigning them silently. Use these definitions consistently across every surface.
| Outcome | Counts when | Does not establish |
|---|---|---|
| Brand mention | The answer identifies the intended company | Positive endorsement or a source link |
| Domain citation | A visible supporting source links to the company's domain | That the company is recommended |
| Commercial recommendation | The company is explicitly proposed for the buying need | First place or exclusive preference |
| Product accuracy | A factual claim matches a verified current product source | That all other claims are correct |
| Technical readiness | Defined checks of site access or content pass | Observed assistant visibility |
Publish the denominator beside every rate
Domain citation rate equals valid observations with at least one qualifying domain citation divided by valid observations in the stated cohort. Count an answer once even when it links to several pages on that domain. Use a separate distinct-URL table to explore which pages receive citations.
Commercial recommendation rate equals valid commercial observations recommending the company divided by all valid commercial observations in that cohort. A negative mention, incidental reference or informational citation is not a commercial recommendation.
For AI Overviews, a valid search without an Overview remains in the overall denominator with no Overview citation or recommendation. Also report trigger rate and conditional outcomes among triggered Overviews. An execution error is different and remains outside these rates.
Report n/N, not only percentages. A top-three measure applies only to explicitly ordered recommendation lists, with its eligible-list count shown. Do not invent a rank for unranked prose.
Work through one illustrative calculation
The following is a fictional arithmetic example, not an Arrow result or an industry benchmark. Assume one commercial question cohort, one public interface and a frozen panel. There are 30 planned attempts: 27 completed answers and three technical failures. The completed answers contain nine brand mentions, six domain citations and four explicit recommendations. These outcomes can overlap.
The mention rate is 9/27, the citation rate is 6/27 and the recommendation rate is 4/27. Three failures are reported as 3/30 attempts; they are not treated as valid non-mentions. Counting several links in one answer does not create additional cited answers. If there were no valid answers, the outcome rates would be not measured.
This small panel describes its collected answers. It does not estimate all users, prove improvement or identify a causal effect from a content change. Add the question coverage and session breakdown before comparing another period.
| Observed outcome (fictional) | Count / eligible total | Interpretation |
|---|---|---|
| Brand mention | 9 / 27 valid answers | 33.3% of the declared sample |
| Domain citation | 6 / 27 valid answers | 22.2%; count each cited answer once |
| Commercial recommendation | 4 / 27 valid commercial answers | 14.8%; identity checks excluded |
| Technical failure | 3 / 30 attempted runs | 10% failed; preserved outside outcome rates |
Save evidence before summarizing it
Keep the complete answer, visible sources, observation ID and classification together. Record an actual shared-result link or capture when available. Avoid replacing the evidence with a generated summary that cannot be checked.
Mark whether a source was visibly cited. Do not claim a page was retrieved merely because it seems relevant: retrieval is an additional event that may not be exposed. Record it only when the collection system provides that evidence.
Review ambiguous brand matches, recommendations and factual claims manually. Recheck a sample of ordinary classifications and document corrections. Store the original classification alongside its revision so that a chart can be regenerated.
- Required record: question ID, panel version, exact prompt, surface, timestamp and session ID.
- Context: language, known market settings, account/session type and collection method.
- Evidence: full output, visible source URLs, completion status and capture reference.
- Review: outcome labels, ambiguity notes, reviewer and revision date.
Use official reports for the outcomes they expose
Google's Generative AI performance report shows impressions in covered Search features, with page, country, date and device views. It does not directly report whether a brand was recommended. Search Console documentation.
Bing AI Performance reports citations across its supported experiences. Keep that platform-provided evidence separate from your custom prompt panel. Bing documentation.
Link answer observations to attributable site visits when evidence permits, then follow audits, qualified inquiries and customers. Retain self-reported discovery information. Avoid assigning every direct visit or subsequent sale to AI.
Make the next measurement comparable
Maintain a page-change log and keep the original question cohort available. Compare question-level outcomes across sessions before interpreting an aggregate increase. A change in answer behavior can have several causes beyond your intervention.
Use the benchmark protocol for a comparison study and the 90-day review guide for decisions. Start with a free audit to identify which questions and site checks are relevant.
Use the downloadable observation kit
Version 2 contains 60 non-branded questions about GEO and AI visibility solutions: 40 commercial and 20 informational questions, plus ten separate Arrow identity controls. Use the commercial questions to assess solution discovery in the stated contexts; exclude the identity controls from that recommendation rate. The questions are a research starting point, not verified search demand.
Download the version 2 question panel, the observation contract and the action journal contract. The kit instructions explain evidence requirements, missing outcomes, repeat sessions and comparable before-and-after periods.
The provided observation journal and action journal are empty. Collect and review real answers before producing a baseline. Store completed records privately; do not publish customer or session data as part of a public template.
The earlier version 1 panel remains available for separate downstream research, including yacht-brokerage and DPE-provider questions. Those subjects and their recommendations must not be merged into the version 2 GEO-solution discovery denominator.
Choose the next measurement decision
Use this page as the common definition of evidence and outcome rates. Each guide below owns a different decision; the measurement collection groups all fourteen guides into four reading paths.
| Your next task | Focused guide |
|---|---|
| Select questions | Prompt selection and panel versions |
| Record the before-state | Baseline before changing content |
| Understand a platform limitation | What ChatGPT visibility can actually measure |
| Choose formulas | The twelve-metric dictionary |
| Build the reporting view | Dashboard layout and data checks |
| Compare companies or periods | Benchmark protocol |
| Prioritize a finding | Sixteen signals and their next actions |
| Run the weekly review | Citation review playbook |
| Make a funding decision | CFO report and decision memo |
| Review healthcare claims | Healthcare citation review |
| Trace a referral | Attribution evidence stack |
| Qualify the inquiry | From buyer question to qualified lead |
| Publish a supported case | Case-study evidence framework |
Sources and editorial scope
This is Arrow AI's implementation guidance. Examples are illustrative unless identified as dated observations. Source access and good content do not guarantee a recommendation.
Continue through the GEO evidence library, inspect Arrow GEO's measurement limits, or start an audit.
