A useful AI visibility benchmark compares companies on the same buyer questions, search surfaces, markets and observation periods. It distinguishes citations from commercial recommendations and shows the underlying counts. This guide provides a proposed research protocol and blank reporting tables; it does not present measured market results.
Choose the business decision the benchmark will support
Decide whether you are evaluating product discovery, category selection or the usefulness of published research. Those questions require different samples. A company can earn citations for an educational report while another company receives the recommendation to buy. Combining both into one leaderboard conceals the difference.
Define one category and a realistic buyer. For example, investigate software that helps a small B2B marketing team inspect AI search visibility. Record the inclusion criteria for the comparison set and explain why each company qualifies. Avoid comparing a broad enterprise suite with a narrow product as if they serve identical needs.
Write the decision before collecting responses: identify missing selection evidence, check product accuracy or prioritize a market. This makes the eventual analysis actionable.
MMention rate
Valid answers naming the subject ÷ valid answers in the declared cohort.
CCitation rate
Valid answers linking to the subject domain ÷ valid answers in the declared cohort.
RRecommendation rate
Valid commercial answers explicitly selecting the subject for the buyer need ÷ valid commercial answers in the declared cohort.
AIdentity accuracy
Valid identity-control answers with a reviewed correct identity ÷ valid identity-control answers with an identity judgment.
Keep the four values separate. A single composite score can hide the difference between being named, being used as a source, being selected for a buyer and being described correctly. If a team creates an internal weighted score, publish the weights and the four source rates beside it.
Build a fixed question panel
Create questions from customer interviews, sales objections and existing search demand. Include discovery, requirements, evaluation and implementation needs. Keep questions that name your company in a separate identity check. Asking an assistant to recommend your company measures a different behavior from unprompted selection.
A proposed pilot contains 20 commercial questions and 10 informational questions, tested on three surfaces across three sessions: 270 planned observations. These numbers are a workload example. Choose a panel your team can repeat and review; they are not a statistically established minimum.
Store each question with an ID, intent, language, target market and reason for inclusion. Freeze the initial version. Add new questions as a separate cohort so that a changing sample cannot silently create apparent growth.
Arrow publishes a versioned proposed question panel and an observation schema so another analyst can inspect the expected fields. The panel is explicitly marked unmeasured; publishing a protocol is not the same as publishing results.
Collect comparable responses
Use a consistent session setup and spread repetitions across several days. Record the exact surface: ChatGPT Search, a particular API configuration and Google AI Mode are distinct environments. Keep separate results if collection methods differ.
Save the full response, visible source links, timestamp, known location settings, language and completion status. A missing response due to a technical error is not the same as a valid response that excludes a company.
For AI Overviews, record whether an Overview appears. In the overall recommendation rate, a valid search without an Overview contributes an observation without a recommendation. Also report a conditional rate among searches where an Overview appeared.
Use definitions a reviewer can apply
The proposed benchmark has three primary outcomes. A mention identifies the correct company. A citation links to its website as a source. A recommendation explicitly presents that company as a solution to the stated buying need. Review homonyms, critical mentions and source-only references before classification.
A citation rate counts responses containing at least one qualifying company-domain citation, divided by valid observations in the stated cohort. Repeating a URL within an answer does not increase this rate. A separate URL-frequency table can count distinct sources.
Commercial recommendation rate uses valid commercial observations as its denominator. Do not include informational questions or identity checks. Record how often the assistant gave any eligible recommendation, which provides context for low category-wide rates. See the full measurement method.
Publish blank fields honestly until collection is complete
The following is a reporting template, not a dataset. Replace pending fields only after responses have been collected and reviewed. Include the question-panel version, period, market and surface above every completed table.
| Company | Valid commercial observations | Responses recommending company | Recommendation rate | Source responses |
|---|---|---|---|---|
| Your company | Pending collection | Not measured | Not measured | Pending collection |
| Comparison company A | Pending collection | Not measured | Not measured | Pending collection |
| Comparison company B | Pending collection | Not measured | Not measured | Pending collection |
Interpret patterns without inventing causality
Compare outcomes by question before comparing overall percentages. A result driven by one question repeated several times is different from gains across multiple buying needs. Show n/N and the spread between sessions. Repeated responses to the same question are not independent people in the market.
Observe which sources support accurate or inaccurate product descriptions. If a competitor has a useful implementation guide and yours does not, that is a testable content opportunity. It is not proof that copying the guide will produce the same result.
Use platform reporting as complementary evidence. Google's Generative AI performance report provides impressions for covered search features; it is not a brand recommendation report. Official Search Console report documentation.
Turn the benchmark into an improvement cycle
Choose one intervention for a defined cohort: clarify pricing, publish a worked example or substantiate a product claim. Keep a change log, repeat the panel and compare the same questions. External updates and competitor activity remain possible explanations for movement.
Publish an exploratory study with its method, limitations and negative results when the collection is ready. Until then, use this protocol for planning. Review Arrow's platform information or request a free audit to scope the questions and evidence relevant to your business.
Use the metric dictionary for formulas and the dashboard guide to display coverage and missingness. The measurement collection connects the full process.
Sources and editorial scope
This is Arrow AI's implementation guidance. Examples are illustrative unless identified as dated observations. Source access and good content do not guarantee a recommendation.
Continue through the GEO evidence library, inspect Arrow GEO's measurement limits, or start an audit.