Use a four-part evidence model
Direct page and crawl facts
Status codes, metadata, extracted claims, sources, dates, headings, canonical URLs and page inventory.
Timestamped provider results
Exact query, model, response time, returned citations, target position and competing domains.
Disclosed comparative estimates
Scores derived from observable signal groups and published weighting, never presented as vendor rankings.
Private system facts
Undisclosed training inclusion, internal weights, unseen user personalisation and guaranteed future results.
Define citation rate before reporting it
A citation rate needs a fixed denominator. For example: target cited in 8 of 40 provider-query observations during a stated seven-day test window. Report the providers, locations, model versions where available, query set, and repeat count. A percentage without this context invites false comparison.
Design a balanced query set
Cover branded questions, category discovery, comparisons, evidence-seeking questions, definitions, implementation questions, and high-value long-tail decisions. Avoid changing the prompt after every unfavourable result. Version the set and rerun it on a schedule so changes are interpretable.
Record source competition
Count the domains cited across the same observations and classify why they may be preferred: primary data, official status, freshness, directness, format, or topic depth. “Competitor X appeared in 18 observations” is more actionable than “we have low AI authority”.
Keep readiness separate from visibility
A page can be citation-ready but not yet observed, and it can be cited despite weak structure. Readiness measures the content and access conditions a team controls. Visibility records what a provider returned. Present both, but do not silently merge them.
Connect observations to business outcomes
Use tagged landing pages, referral reporting where available, branded search lift, assisted conversions, sales-call mentions, and content engagement. Attribution will be incomplete, so label direct, assisted, and qualitative evidence separately.
A defensible report line
“IndexHalo observed the target domain in 6 of 30 repeatable provider-query checks between 1 and 7 August. The result is an observed citation rate for this test set, not evidence of model training or a prediction of future citations.”
Sampling, variance and the confidence interval
Answer engines are non-deterministic. Ask the same question twice and you can get different sources, different phrasing, and a different set of cited domains. Any measurement framework that ignores this will report noise as progress.
The practical consequence is that a single observation is not a measurement. If you query once and see a citation, you have learned that the citation is possible, not that it is likely. Presence rate — the proportion of runs in which you appear, across repeated samples of a fixed query set — is the smallest honest unit, and it needs a confidence interval attached.
Interval width is driven by sample size. Ten runs of twenty queries produces intervals wide enough that most week-to-week movement is indistinguishable from chance. This is the uncomfortable arithmetic behind most GEO dashboards: they display a number to one decimal place that their sampling cannot support. If a change falls inside the interval, the correct report is “no detectable change”, not a percentage.
Two disciplines keep this honest. Fix the query set and change it only deliberately, with a version number, because altering the questions between runs means you are measuring the questions. And report the interval next to the point estimate every time, so a reader can see whether a movement is real before acting on it.
Connecting citations to business outcomes
The hardest question in this discipline is not whether you are cited but whether it matters, and the honest answer is that the attribution chain is weak and should be described that way.
Some AI surfaces pass a referrer, and those visits appear in analytics as identifiable referral traffic. Many do not. A user who reads an answer, forms an impression, and searches your brand name a week later is indistinguishable from organic brand demand. You cannot close that loop with tracking, and vendors claiming to have done so are inferring, not observing.
What you can do is triangulate. Watch direct referral traffic from AI surfaces where it exists. Watch branded search volume against periods where citation share moved. Ask new customers, in a form field, where they first heard of you — self-reported attribution is noisy but it is real data, and it consistently surfaces AI surfaces earlier than analytics does.
Then report the chain with its uncertainty intact: citation share is measured, referral traffic is measured, and the causal link between them is inferred. A report that presents the inference as a measurement will eventually be checked by someone, and the whole programme loses credibility at that point.
Designing a query set you can defend
Everything downstream depends on the query set, and most are assembled badly — usually as a list of keywords lifted from an SEO tool and punctuated with a question mark.
Answer surfaces receive questions, not keywords, and the phrasing carries intent that a keyword strips out. “CRM pricing” and “how much should a mid-market CRM cost for 200 seats” retrieve different sources, and only the second resembles what people actually type. Build the set from real language: support tickets, sales-call transcripts, the questions prospects ask in demos. The engine-specific guides for ChatGPT, Perplexity and AI Overviews each describe how queries reach that surface.
Cover the funnel deliberately rather than clustering at the bottom. Include definitional questions where you would expect to be cited as a reference, comparative questions where competitors will appear, and decision questions with commercial intent. A set weighted entirely to the last group will report that you are invisible, because those queries are the most contested and the least representative of where citation is winnable.
Then freeze it and version it. Every change to the wording changes what you are measuring, so amendments belong in a new version with a documented reason, and trend comparisons should only ever run within a version. A query set quietly edited between runs produces a chart that looks like progress and means nothing at all.
Add governance to the dashboard
- Show the evidence class beside every metric.
- Preserve timestamps and raw returned citation URLs.
- Version query sets and scoring weights.
- Do not replace missing data with a confident-looking zero.
- Keep sample or demonstration data visibly labelled.
- Document provider outages, rate limits, and unconfigured checks.
Primary references
Common questions
Can anyone actually measure AI visibility?+
You can measure specific, bounded things: whether a URL was returned as a citation for a stated query, at a stated time, from a stated provider. You cannot measure what a model was trained on, how it ranks sources internally, or whether it will cite you tomorrow. A credible report is explicit about which of those it is doing.
How many queries do I need for a reliable baseline?+
More than most dashboards use. Twenty queries sampled a handful of times produces confidence intervals wide enough that most week-to-week movement is indistinguishable from noise. Either widen the query set, increase repeat sampling, or report changes as undetectable until they clear the interval.
What is a good presence rate?+
There is no universal benchmark, and any vendor quoting one is inventing it. Presence rate is only meaningful against your own baseline on a fixed query set, or against named competitors on the same set. The trend across dated runs is the signal; the absolute number is not.
Why do vendors report such different numbers?+
Because they use different query sets, different sampling rates, different providers and different definitions of a citation. Two tools can both be honest and disagree substantially. Ask what the query set is, how many samples underlie each figure, and whether the raw archived answers are available.
Should I trust a tool that will not show raw responses?+
No. An aggregate without the underlying observations cannot be checked, and an unverifiable number is decoration. Any credible platform, including ours, should let you drill from a headline metric back to the exact archived answer, query, provider and timestamp that produced it.
How do I connect citations to revenue?+
Carefully, and with the uncertainty stated. Track referral traffic from AI surfaces where it exists, watch branded search against periods where citation share moved, and ask new customers where they first heard of you. Report the citation data as measured and the revenue link as inferred — presenting the inference as measurement is how programmes lose credibility.