Use separate measures for collection quality, observed answers and business outcomes. Every rate needs its numerator, denominator, period and sample definition. Missing evidence is not zero; a citation is not automatically a recommendation, and neither proves a sale.
Start with a metric register, not an average score
A metric is useful only when another analyst can reproduce it. Before calculating anything, record the evaluated company and domain, prompt-panel version, language, intended market, interface, model when disclosed, session context and observation dates. Report public interfaces separately from provider APIs. A response returned by an API does not describe what every user sees in a consumer product.
The definitions below are an operational scorecard, not industry benchmarks or a published ranking formula. Use the citation tracking method for collection rules. Label each result with its counts; retain the underlying answer, citation URLs and review decision. If an eligible denominator is empty, show “Not measured.”
Metrics 1–2: can the sample support a comparison?
Collection quality comes first. A percentage from three successful answers should not look as complete as one from a fully observed panel. Keep failed attempts visible and prevent retries from becoming extra independent successes.
For prompt coverage, choose the eligible panel group before collecting. Ten identity questions and sixty non-branded questions answer different questions and must not share a coverage denominator.
| Metric | Calculation | Interpretation |
|---|---|---|
| 1. Valid-observation rate | Valid recorded observations / recorded attempts, with retries identified | Collection completion; errors and unavailable outputs are not brand absences. |
| 2. Distinct-question coverage | Distinct questions with at least one valid observation / eligible questions in the frozen group | Breadth within your chosen panel; not coverage of all buyer demand. |
Metrics 3–6: distinguish presence, sourcing and selection
A confirmed mention identifies the intended company, not an unrelated business with the same name. A source citation is an observed supporting link to the subject domain. A recommendation explicitly proposes the company for the buying need. Review each independently.
For a valid Google search with no AI Overview, preserve that outcome in overall search-observation rates and separately report whether an Overview appeared. A rate conditional on triggered Overviews needs a different, clearly named denominator.
| Metric | Numerator / denominator | Review rule |
|---|---|---|
| 3. Confirmed brand mention rate | Valid observations mentioning the intended company / eligible valid observations | Resolve identity using the domain and captured context. |
| 4. Subject-domain citation rate | Valid observations with an observed subject-domain supporting link / eligible valid observations | Count each observation once, even if it cites several subject URLs. |
| 5. Commercial recommendation rate | Valid non-branded commercial observations explicitly recommending the subject / valid non-branded commercial observations | Exclude identity controls and informational questions. |
| 6. Identity accuracy rate | Correct identity-control answers / reviewed valid identity-control answers | Use declared reference facts; disclose the proportion actually reviewed. |
Metrics 7–9: expose review gaps and answer variation
Fact accuracy applies to checked claims, not entire answers. Archive the applicable pricing, feature or policy reference with its date and version. Classify unsupported claims as unverifiable rather than silently treating them as correct or incorrect.
Repeat agreement describes stability, not truth. Compare the same questions, contexts and classification rule across independent sessions. A brand can be consistently absent or consistently misdescribed. Do not interpret high agreement as high visibility.
| Metric | Calculation | Required companion information |
|---|---|---|
| 7. Fact-review coverage | Valid answers reviewed for eligible subject claims / eligible valid answers | Reviewed answers with no eligible claims remain reviewed; missing review remains unknown. |
| 8. Verified fact accuracy | Correct checked claims / (correct + incorrect checked claims) | Show unverifiable claim counts and category breakdown separately. |
| 9. Classification repeat agreement | Comparable question-context pairs with the same reviewed result / comparable reviewed pairs | Name the classification, such as recommended versus not recommended, and report unpaired records. |
Metrics 10–12: measure observed visits and inquiry quality
These measures come from analytics and qualification records, not from a prompt panel. OpenAI documents ChatGPT referral tagging; this helps identify some inbound traffic but does not reveal every prior AI interaction or the visitor’s original prompt.
Use a documented source rule and a consistent period. In GA4 Traffic acquisition, session-scoped dimensions support a session-level view. Keep self-reported discovery and separately attributed CRM activity alongside it, without adding overlapping counts.
| Metric | Calculation | Boundary |
|---|---|---|
| 10. Observed AI-source sessions | Count of sessions matching the declared AI-source rule in the period | A count has no rate denominator. Preserve the source list and analytics coverage. |
| 11. AI-source key-event session rate | Matching AI-source sessions with the chosen key event / matching AI-source sessions | Choose one defined event, such as a submitted demo request; count sessions rather than repeated events. |
| 12. Qualified AI-linked inquiry rate | Reviewed AI-linked inquiries meeting the qualification rule / reviewed AI-linked inquiries | Deduplicate inquiries, disclose unreviewed inquiries and define how the AI link was established. |
A worked example: three rates from the same small panel
Illustrative example, not Arrow or client results: a team records 20 independent attempts for ten non-branded commercial questions. Eighteen return valid answers; two fail. Eight valid answers mention the company, five include a subject-domain citation and three explicitly recommend it. All ten questions have at least one valid answer.
The valid-observation rate is 18/20, or 90%; question coverage is 10/10, or 100%. Mention rate is 8/18, citation rate 5/18 and recommendation rate 3/18. Dividing by 20 would incorrectly classify two collection failures as negative answers. Adding the three answer counts would double-count overlapping observations.
No analytics or CRM evidence was collected in this example. Metrics 10–12 remain unmeasured. Publish the counts and the small sample limitation before deciding whether another observation cycle or a source-page correction is warranted.
Sources and editorial scope
This is Arrow AI's implementation guidance. Examples are illustrative unless identified as dated observations. Source access and good content do not guarantee a recommendation.
Continue through the GEO evidence library, inspect Arrow GEO's measurement limits, or start an audit.
