Start with the denominator
Define the target brand, real competitor domains, a fixed portfolio of commercially important buyer questions and a fixed provider set before calculating any share. Twenty-four questions checked across three providers create seventy-two attempted responses. If only sixty-eight responses succeed, the evidence base is sixty-eight verified responses—not seventy-two assumed negatives.
The resulting share describes only that configured answer market. It is not universal market share, prompt volume, audience reach, training inclusion or a provider’s private preference. Changing the brand, question or provider universe creates a different benchmark. Preserve the exact universe and fingerprint so later movement is defensible.
Resolve brands with exact disclosed identities
Each brand needs a canonical domain and a short list of exact aliases. The target aliases can include the monitored organisation, approved product names and domain. Competitors come from domains the customer deliberately configures and public titles measured during the competitor crawl. Weak terms such as “home” or “official” must be rejected because they create false matches.
Preserve the original Unicode response. A Chinese, Arabic, Hindi, Japanese or Korean brand name should be matched as published rather than converted into an English-only token. For spaced scripts, compare bounded normalized phrases so a short brand does not accidentally match part of another word.
Separate four answer-market signals
- Mention share. How often each configured brand is named in verified response text, divided by all configured-brand mentions.
- Recommendation share. How often a sentence both names the brand and contains disclosed recommendation language, divided by all such response labels.
- Citation share. How often an exact configured brand domain appears in returned response sources, divided by all configured-brand citation appearances.
- First-choice share. Among responses with recommendation language, which brand appears first in that recommendation context.
These signals answer different questions. A brand can be frequently cited as a data source yet rarely recommended as the product to choose. It can be mentioned in a comparison while a competitor is placed first. Reporting one blended “visibility” number without the four components hides the work a team actually needs to do.
Use a transparent market index
IndexHalo’s disclosed index weights mention share at 35%, recommendation share at 35%, citation share at 20% and first-choice share at 10%. If a signal has no observations anywhere in the configured set, its weight is removed and the remaining weights are normalized. This avoids awarding or penalising brands for evidence that did not exist.
The leaderboard must show the component shares, counts, verified checks, rank and observed lexical stance beside the index. A one-brand universe is withheld because it would mechanically return 100%. A run with no verified response text is labelled not observed. Neither case should be filled with an estimate.
A mention is not an endorsement. Recommendation is a deterministic multilingual review label, not a provider-supplied intent field. Returned citations are response-level unless the provider exposes durable statement-level attribution.
Turn the leaderboard into question battles
Portfolio share tells executives where the brand stands; question-level evidence tells teams what to do. For every question, retain the exact leader, target rank, index gap, mentions, recommendations, citations, first choices, provider count and mapped owning URL. Classify a material target lead, a narrow contested result, a competitor lead, target absence or a completely unobserved market.
If a competitor leads a German warehouse payback question, compare the mapped owner’s public evidence with the winning public page. The action might specify building size, tariff period, financing cases, sample, source proximity and decision thresholds. If the target already leads, protect its canonical URL, method, dataset and entity naming instead of rewriting successful evidence.
Inspect provider differences without inventing causes
Calculate the same brand rows separately for every successfully observed provider. A target may lead citations in one provider but trail recommendations in another. That difference is operationally useful, but it does not reveal private ranking logic. Report the provider, verified-response count, target rank, target index, leader and measured gap.
Do not fill a failed provider response with zero, and do not extrapolate three monitored providers to every LLM product. The purpose of provider segmentation is to inspect reproducible observations and plan another equivalent check—not to make population-wide claims.
Compare only like with like
Period movement requires the same question-set fingerprint, successful provider set and configured brand universe. Preserve excluded runs with a reason such as changed questions, changed providers or no verified response text. A ten-point gain after doubling the number of branded questions is not comparable progress.
A controlled experiment can freeze the target question, mapped page, market fingerprint, entity domains, provider set, verified-response count, baseline target index, rank and leader gap. After the change ships, the later run passes only when those guardrails remain comparable and the saved improvement threshold is met.
Convert share gaps into accountable work
Every action should contain the current leader and score, target score and rank, exact owner, recommended production change, accountable team and definition of done. For a target-absent question, build one evidence-backed owner. For a material competitor lead, strengthen the existing owner. For a contested result, add defensible differentiation. For a target lead, preserve the winning evidence bundle.
The verification should repeat the unchanged question across the same providers and compare exact aliases, recommendation sentences, citation domains and first-choice order. A practical target can require the target to rank first, receive recommendation labels in at least two verified responses and lead the next configured brand by ten index points.
Operate across commercially important languages
Recommendation classification should recognize decision language beyond English and preserve excerpts for human review. Local-language monitoring must keep original questions, brand spellings, citations and evidence owners. Translation alone is insufficient: regulation, tax, pricing, units, sources and buyer criteria must be locally correct.
A multilingual benchmark can compare the same strategic theme across markets, but each language and jurisdiction remains its own evidence context. A Spanish incentive owner should cite the current administering authority; a German payback page should expose German energy assumptions. The goal is equivalent decision usefulness, not identical copy.
The bottom line
AI share of voice becomes a professional management tool when every percentage can be traced to a stored answer, exact brand rule, returned domain and fixed denominator. The valuable output is not the headline rank alone. It is the question-by-question decision, production specification and repeatable test that tells a team whether its next release actually won more of the configured answer market.