IndexHalo
Strategy

How to Choose a GEO Tool

The GEO tooling category went from nearly empty to crowded in about eighteen months, and the marketing has outrun the methodology. This is a buyer's framework: what to test during a trial, which questions expose weak measurement, and the cases where the honest answer is that you do not need a tool yet.

IndexHalo Editorial Team9 min read

Start by asking whether you need one

The most useful thing a buyer's guide can say first is that the entry-level version of this work needs no software at all.

Write down twenty questions a prospective customer would genuinely type. Ask each one in ChatGPT, Perplexity and Google, on a fixed cadence, and record in a spreadsheet whether you were cited, which URL was returned, which competitors appeared, and the date. Paste the full answer text into a column. That is a real baseline, it costs about an hour a fortnight, and it will tell you more than most dashboards.

Tooling becomes worth paying for at three thresholds: when the query set grows past what a person will reliably collect by hand, when you are tracking several brands or competitors and the combinatorics get out of hand, or when you need archived, timestamped evidence you can hand to someone else. Below those thresholds, a tool mostly automates a spreadsheet you have not yet proved you will maintain.

Four things to test in a trial

Trials are usually spent exploring features. That is the wrong use of them. Four tests tell you more than the whole feature list.

Can you reach the raw answer? Pick any headline number and try to drill from it back to the exact archived response, with its query, engine, model and timestamp. If you cannot, the number cannot be checked, and an unverifiable number cannot support a decision. This single test eliminates a surprising share of the category.

Is it stable? Run the same query set twice, a day apart, changing nothing. Some movement is expected — these systems are non-deterministic. Large unexplained swings mean the sampling is too thin to detect the changes you will actually care about.

Does it show uncertainty? Look for sample sizes and confidence intervals next to the percentages. A tool reporting presence rate to one decimal place off five samples is presenting noise with false precision, and you will end up explaining meaningless movements to stakeholders.

Does it separate measurement from inference? Crawl findings are measured. Provider responses are observed at a timestamp. Comparative estimates are modelled. Training data and internal rankings are unknown. A tool that labels these distinctly is one you can defend in a meeting; a tool that presents them uniformly will eventually embarrass you.

Engine coverage, and why more is not automatically better

Coverage counts are a favourite marketing number and a weak decision criterion. What matters is coverage of the surfaces where your buyers actually are.

A B2B software company whose buyers research in ChatGPT and Google gains little from a tool that also polls three engines with negligible share in its market, and may pay per query for the privilege. Conversely, a consumer brand may find a surface irrelevant to B2B is where most of its exposure sits.

Ask how each engine is accessed, because the answer affects both cost and fidelity. API access is metered and reproducible. Scraping a consumer interface is neither reliably permitted nor stable. Manual observation — pasting in answers a human saw — is free and honest but does not scale. Tools differ substantially here and rarely lead with it.

Reading the pricing model

Two structures dominate, and each hides its costs in a different place.

Query-metered pricing is transparent about marginal cost and suits a focused query set monitored often. The failure mode is a large prompt set multiplied across several engines and repeat samples, producing a bill nobody previewed. Ask whether the tool shows you the cost of a sweep before it runs, and whether it warns about near-duplicate prompts that consume calls without adding information.

Seat pricing suits agencies and larger teams, and its failure mode is the opposite: unclear limits on the underlying volume, discovered when a crawl silently truncates or a sweep is capped. Ask what the actual ceilings are and what happens at them — a bounded crawl that reports what it excluded is fine; one that silently stops is not.

The questions that expose weak methodology

Ask these in a demo. The quality of the answers is more informative than any feature comparison.

How many samples underlie a reported presence rate, and is the interval shown? What exactly counts as an appearance — a linked citation, a brand mention in the text, or both, and can they be separated? What happens when a query returns no answer at all: does that count as an absence or is it excluded? How is the query set versioned when it changes, and can trends still be compared across a version boundary? What does the tool explicitly claim not to know?

A vendor with rigorous methodology will enjoy these questions, because they are the parts they got right. A vendor without one will redirect to the roadmap.

Our own position, stated plainly

We build one of these tools, so treat this section accordingly and apply the tests above to us as readily as to anyone else.

What we think we do well: every metric drills back to the raw archived answer; presence rate is reported with its sample size and confidence interval; findings are labelled as measured, observed, modelled or unknown; and sweep costs are previewed before they run. Those are the four tests above, and we designed around them because we kept encountering reports that failed them.

What we do not claim: any knowledge of training corpora, any insight into internal ranking weights, and any ability to guarantee a future citation. If a comparison matters to you, run the same query set through two tools for a fortnight and compare what each will show you underneath the headline number. That is a better basis for a decision than either vendor's description of itself.

Common questions

Do I need a GEO tool at all?+

Not to start. A spreadsheet of twenty questions, asked manually across engines on a fixed fortnightly cadence, produces a real baseline for the cost of an hour. Tooling earns its place when the query set or the number of brands makes manual collection impractical, or when you need archived evidence for reporting.

What should a trial actually test?+

Whether you can reach a raw archived answer from any headline number, whether the same query set produces stable results across runs, whether confidence intervals are shown, and whether the tool distinguishes measurement from estimation. Feature lists are far less informative than these four.

Why do two tools report different numbers for my brand?+

Different query sets, sample sizes, engine coverage and definitions of an appearance. Both can be internally honest and still disagree substantially. This is why comparing absolute scores between tools is meaningless and comparing trends within one tool is not.

Is per-seat or per-query pricing better?+

Depends on where your cost actually sits. Query-metered pricing suits a small, high-value query set monitored frequently. Seat pricing suits agencies running many brands. The trap in both is unmetered sweeps, where a large prompt set across several engines produces a bill nobody previewed.

How much should I expect to pay?+

The category spans free tiers through to enterprise contracts, and price correlates weakly with methodological rigour. A cheap tool that archives raw answers is more useful than an expensive one that reports a composite score you cannot audit.

What is the biggest red flag?+

A vendor who will not show you the raw response behind a metric. Everything else — thin engine coverage, an odd pricing model, a sparse feature set — is a trade-off you can evaluate. An unverifiable number is not a trade-off; it is a number you cannot use.