Methodology · Corpus v1.0 (frozen 2026-08-11)

How the Agent Visibility Index is measured

Every number we publish is designed to survive scrutiny. This page describes the full pipeline; it is the contract behind every score.

What we measure

How often AI assistants recommend each brand when asked realistic buyer questions — the Share of Recommendation — plus first-mention rate, sentiment (clean / caveat / recommended-against), and the reasons assistants give for choosing or rejecting each brand.

The task corpus

Sampling

Each task runs per model surface per weekly cycle (models are stochastic; single answers are anecdotes). Raw responses are stored verbatim, forever — every historical score can be re-derived and re-audited.

Statistics

Model-update detection

A ~24-task control set of stable scenarios runs every cycle. If its brand distribution shifts beyond a total-variation threshold, the cycle is flagged and the surface's epoch increments. Deltas across an epoch boundary are annotated "model update" and are never labelled significant — you should never mistake a model change for a market change.

What "the model said" means

We query each provider's API with an identical, neutral system prompt across all surfaces. This is a controlled proxy for consumer chat products, which layer their own prompts, memory, and retrieval on the same models. It maximizes comparability across surfaces; it is not a screenshot of any one app. Grounded (web-search-enabled) surfaces are tracked separately as they ship.

Measured surfaces

SurfaceidPinned modelModeStatusEpoch
Claude (fast) anthropic-fast claude-haiku-4-5-20251001 plainactive1
ChatGPT (API) openai-chat gpt-5.2-2025-12-11 plainpaused1

Model versions are pinned, never aliases — epochs rotate only at cycle boundaries, and the full history is below.

SurfaceEpochModelFrom
anthropic-fast1 claude-haiku-4-5-202510012026-08-11
openai-chat1 gpt-5.2-2025-12-112026-08-11

The judge

Structured extraction (which brands, what position, what sentiment, what reasons) is performed by an LLM judge with a strict JSON schema and alias normalization against the corpus brand dictionary. The judge-agreement rate on a 200-response hand-labelled sample is pending publication — it will appear here before the public index launches.

Neutrality commitment. Rankings are never for sale, at any price. We sell analytics about the rankings — never positions on them — and measurement is never bundled with paid "improve your score" services.