# Reproducible AI observation kit

This kit contains a proposed question panel, schemas and empty journals. It contains **no collected baseline, vendor ranking or measured market result**. Synthetic responses exist only in repository tests.

## Downloads

- [Primary Arrow question panel v2](./prompt-panel.v2.json)
- [Primary observation schema](./observation.schema.json)
- [Frozen secondary supplier panel v1](./prompt-panel.v1.json)
- [Archived v1 observation schema](./observation.v1.schema.json)
- [Empty observation journal](./observations.empty.jsonl)
- [Action journal schema](./action-log.schema.json)
- [Empty action journal](./actions.empty.jsonl)
- [Blank monthly decision memo](./monthly-decision-memo.md)
- [Header-only attribution ledger](./attribution-ledger.csv)

## Panel scope

The primary v2 panel has 60 non-branded questions about GEO solutions: 40 commercial and 20 informational. Ten additional questions check Arrow AI's identity. English and French each have 30 non-branded questions and five identity controls. Arrow is the primary evaluated subject; the questions do not name Arrow outside identity controls.

All three non-branded cohorts have 20 questions and subject kind **geo_platform**. The **b2b** cohort concerns general B2B GEO purchasing and evaluation. **yachts** asks about GEO solutions for a yachting business; **dpe** asks about GEO solutions for French diagnostic businesses. These are buyers choosing software or implementation services, not consumers choosing brokers or diagnosticians. Every prompt declares language, intended market, intent and eligible subject kind. Records for other GEO vendors remain separate by subject and domain.

The frozen **v1** panel is retained as a secondary supplier study: its yacht/DPE questions evaluate downstream providers. It must not be described as 40 commercial questions evaluating Arrow. Select it explicitly with **--panel geo/authority-library/measurement/prompt-panel.v1.json**, use its archived schema, and keep its runs separate. Its original record shape excludes the new ranking/fact-review fields; v1 drafts retain that shape. No v1 history or answers have been rewritten. Version identifiers, questions and evaluated subjects cannot be combined across v1/v2.

Review these proposed questions for business relevance before collection. They are not verified demand or search-volume data. Freeze the wording for each cycle; changes require a new version.

## Local workflow

The Arrow repository includes a dependency-free CLI in **scripts/geo-loop/**, requiring Node.js 20 or later. It makes no service calls and never infers missing answers.

~~~sh
node scripts/geo-loop/cli.mjs init --dir .drafts/geo-loop/run-first
node scripts/geo-loop/cli.mjs draft --prompt b2b-en-01 --out .drafts/geo-loop/run-first/draft.json --panel .drafts/geo-loop/run-first/panel.json
~~~

Complete the draft only after collecting real evidence. Place reviewed objects on individual lines in an import JSONL file:

~~~sh
node scripts/geo-loop/cli.mjs validate --input reviewed.jsonl --panel .drafts/geo-loop/run-first/panel.json
node scripts/geo-loop/cli.mjs import --input reviewed.jsonl --log .drafts/geo-loop/run-first/observations.jsonl --panel .drafts/geo-loop/run-first/panel.json
node scripts/geo-loop/cli.mjs report --input .drafts/geo-loop/run-first/observations.jsonl --out .drafts/geo-loop/report-first --panel .drafts/geo-loop/run-first/panel.json
~~~

Initialization saves an exact **panel.json** snapshot and its SHA-256 digest in **run.json**. Preserve that snapshot and supply **--panel .drafts/geo-loop/run-first/panel.json** to subsequent commands in the cycle. A new wording requires a new panel version. All commands accept **--panel** for an explicitly selected panel. New drafts/reports never overwrite existing files. The French procedure is in **scripts/geo-loop/OPERATIONS.md**.

The CLI defaults to v2. A proposed baseline is 60 non-branded questions × four public surfaces (ChatGPT Search, Perplexity, Google AI Mode and Google AI Overviews) × three independent sessions: **720 planned attempts**, spread across days. Identity controls are additional and reported separately. This is a collection plan, not a completed sample; actual valid answers, errors and unavailable surfaces determine observed denominators.

## Observation contract

Each JSONL line is one attempt for one evaluated subject. The schema lists required fields; the CLI also checks semantic consistency.

- **id** identifies an attempt. **sampleId** identifies a question/surface/subject/session unit. Retries keep the sample/session; independent repetitions use new sessions. Only one valid outcome is allowed per sample.
- **sessionContext** is fresh_logged_out, fresh_logged_in, existing_conversation, api_stateless, api_stateful or unknown. Reports keep these contexts separate. Prefer comparable fresh sessions; unrecorded conversation history limits comparison.
- **promptId, promptText, panelVersion, language, market** must match the panel exactly. Market states intended audience, not geolocation.
- Record the actual **surface**, known **model** and **collectionMethod**. Public surface names are chatgpt_search, perplexity, google_ai_mode, google_ai_overviews, claude_search, gemini and copilot. APIs require an **api/provider/configuration** surface with collectionMethod **api**; public surfaces require **public_interface**.
- Unknown model: **not_disclosed**. Unknown observed location: **null**, with locationMethod **unavailable**. Other location methods are **account_setting, explicit_ui, other_recorded**.
- **subject** needs ID, name, domain and a kind matching the panel. Arrow identity controls require ID **arrow-ai** and domain **arrow-ai.us**.
- A valid answer needs its complete actual **rawResponse**, real **evidenceRefs**, sources and a named, dated reviewer. The validator cannot authenticate captures or check their accessibility; the reviewer must do that.
- **classification.mentioned** identifies the intended company. **recommended** means explicitly proposed for the buying need and requires a mention. **citationUrls** must be observed supporting links on the subject domain, included in **sourceUrls**. Citation does not imply recommendation.
- **identityCorrect** is a reviewed boolean only for identity controls; otherwise null.
- **error/unavailable** need **errorReason**, with classification, reviewedBy and reviewedAt null. Preserve evidence of the attempt; missing outputs never become zero-valued outcomes.

For a valid Google search without an Overview: **overviewPresent: false**, real search capture, rawResponse null, sourceUrls/citationUrls empty, mentioned/recommended false and identityCorrect null. Triggered Overviews use true. Other surfaces and failed attempts use null.

Use actual UTC ISO timestamps. Keep real completed journals and evidence private in your local workflow; public templates remain empty.

## Optional ordered-list and fact review

**ranking** is null or omitted until reviewed. An object records **explicitlyOrdered, subjectPosition, listLength, evidenceRefs**. Use true only when the answer explicitly ranks or orders a relevant list of GEO providers; display sequence, bullets or prose alone do not establish a ranking. The reviewer must identify the relevant list in the captured evidence. For a genuinely ordered list, listLength is its positive length and subjectPosition is the subject's one-based position, or null if absent from that list. A reviewed answer with no explicit order uses false and null for both position and length. Ranking evidence must reference the captured response evidence. Do not turn unreviewed answers into absent subjects.

Reports limit top-three metrics to non-branded commercial questions and show three denominators: reviewed rankings / valid observations; explicitly ordered lists / reviewed rankings; subject positions 1–3 / explicitly ordered lists. No eligible ordered lists means a null top-three rate. Presence in an ordered list remains distinct from the separately reviewed recommendation flag.

**factChecks** is null or omitted until reviewed; an empty array means the response was reviewed and contained no eligible factual claims about the evaluated subject. Each checked claim is an exact response extract with a unique **id**, **category** (price, feature, category or limit), **verdict** (correct, incorrect or unverifiable), **referenceUrls**, **referenceVersion**, **checkedAt** and real **evidenceRefs**. Use a dated/versioned primary reference and preserve its evidence; record billing period, currency, tier or applicable configuration where relevant. Check every eligible claim using the same rule, not only favorable claims. An unverifiable claim still records the sources checked and why they do not establish its truth in the evidence. The final reviewer verifies attribution to the subject, reference authority and currency, and completeness; the CLI cannot establish those facts.

Fact accuracy is correct claims / (correct + incorrect claims), with separate category counts, unverifiable counts and review coverage. Unverifiable facts do not silently become false or correct. Repeated claims in independent answers remain separate observations, not a count of unique facts across the market. Unknown or absent reviews stay unmeasured. Reference URLs used for fact checking are not AI citations unless actually present in the observed answer's sources. **identityCorrect** remains a separate identity-control measure.

## Reports and comparisons

Reports separate subjects, domains, cohorts, intents, controls, surfaces, declared models, collection methods, languages, intended markets and observed locations. JSON includes daily UTC groups, evidence IDs, mentions/non-mentions and Overview trigger/conditional rates.

Rates expose counts and denominators. Commercial recommendation rate covers only commercial non-branded observations. Identity accuracy uses a separate eligible-answer denominator. Coverage means distinct valid questions in the relevant panel group, not market share or invented planned repetitions.

Valid searches without an Overview remain in overall denominators. Errors/unavailable attempts are reported separately. No valid observations means no measured baseline.

~~~sh
node scripts/geo-loop/cli.mjs compare --before before.jsonl --after after.jsonl --out .drafts/geo-loop/comparison-first --panel .drafts/geo-loop/run-first/panel.json
~~~

Comparison requires successive non-overlapping valid periods and independent observation sessions. Reusing the same successful sample/session with a new record ID is rejected. It pairs matching contexts/questions, giving each question equal weight after averaging its valid sessions. Changed models/methods/observed locations remain unmatched. Unknown settings still limit comparability. Differences are descriptive; no causality or statistical significance is claimed.

Top-three comparisons additionally require explicitly ordered lists for each paired question in both periods. They are conditional on that subset; changes in ordering frequency can change eligibility. Report the ordering and review coverage alongside any change. Fact accuracy is reported per run and is not automatically compared: claims and reference versions may differ, so a reviewer must establish a comparable factual sample first.

## Action journal

Link prompt IDs and before observations to a URL, proposed change, patch references, retest observations and decision. States: **planned, implemented, retested, closed**. Implementation needs a dated patch; retest needs observation IDs; closure needs a decision, which may be inconclusive. Failed retests remain failed attempts.

~~~sh
node scripts/geo-loop/cli.mjs validate-actions --input actions.jsonl --observations all-observations.jsonl --panel .drafts/geo-loop/run-first/panel.json
node --test scripts/geo-loop/loop.test.mjs
~~~

Supply the combined actual observation log for reference, subject and chronology checks.

Every action requires at least one real **beforeObservationId**. A site edit without an AI baseline belongs in the implementation review, not in a fabricated action record. Keep the action journal empty until evidence exists; a later first observation is a baseline for the next cycle, not a reconstructed “before” result for an earlier edit.

Read [Arrow's measurement method](/blog/ai-citation-tracking/) for the editorial definitions. Official reports have their own coverage: [Google Generative AI performance](https://support.google.com/webmasters/answer/16984139) and [Bing AI Performance](https://blogs.bing.com/webmaster/February-2026/Introducing-AI-Performance-in-Bing-Webmaster-Tools-Public-Preview).

## Business evidence templates

The memo and CSV are manual reporting aids, not CLI input formats or automated
connectors. The CSV contains only headers and no customer data. Keep completed
records private and access-controlled. Use consented internal identifiers where
permitted; do not copy names, emails, full query strings or raw customer messages
into a public report. Store masked paths and evidence references where possible.
Blank or unknown values mean not measured; do not infer an AI source from direct
traffic. Self-reported discovery and observed referrals can overlap and must not
be added as unique leads. Preserve the analytics scope and qualification rule
alongside each outcome. A sampled AI answer cannot be joined to an actual buyer
unless independent, authorized evidence establishes the connection.
