Most GEO ROI claims are unprovable. That does not mean the work has no return. It means the measurement has to be built deliberately, and the limits stated out loud.
This page sets out how we measure GEO, what we instrument, what we can only infer, and where we tell clients the honest answer is "we cannot separate this from noise yet". If you are evaluating a vendor, the questions in what to ask before you buy are the fastest way to find out whether they are measuring or guessing.
Why GEO attribution is genuinely hard
Classic search attribution rests on a click. A ranked link is requested, a referrer header arrives, a session begins. Generative answers break that chain in four places.
Answers are often terminal. A buyer who reads a synthesised answer naming three vendors may never click any of them, then search your brand directly a week later. The AI surface created the demand; your branded search gets the credit.
Answers are non-deterministic. The same prompt asked twice can return different sources. Any single observation is a sample, not a measurement. Evertune's approach of sampling each prompt 100 times across models exists precisely because of this variance.
Referrer data is incomplete. Some assistants pass a referrer, some do not, and user-triggered fetches behave differently again from indexing crawls.
There is no rank to track. There is no position one. There is share of an answer that may or may not be generated, phrased differently each time.
What can actually be instrumented
Four things are measurable with reasonable confidence.
| Measure | What it tells you | Confidence |
|---|---|---|
| Share of answer | How often you appear across a fixed prompt set, sampled repeatedly | Good, if the prompt set is fixed and sampling is repeated |
| Citation share | Your share of cited sources versus named competitors | Good, same caveat |
| AI referral sessions | Sessions arriving from assistant surfaces that pass a referrer | Partial: undercounts by design |
| Assisted conversions | Branded search and direct traffic lift correlated with visibility gains | Inferential, not causal |
The first two are the honest core of GEO reporting. They are leading indicators. They do not prove revenue, and any report presenting them as revenue is overclaiming.
The scale reality nobody selling GEO mentions
Per the Reuters Institute's Journalism, Media and Technology Trends and Predictions 2026, Google still delivers roughly 500 times more referrals than ChatGPT from search alone, and around 1,300 times more when Discover is included. The same report notes AI Overviews appear in roughly 10% of US search results, and that Google organic referrals to 2,500+ news sites fell 33% globally and 38% in the US between November 2024 and 2025.
Both facts are true at once. Traditional search is still the overwhelming majority of referral volume, and its share is eroding. That is why we treat GEO as an addition to search work, never a replacement for it. Any vendor telling you to move budget wholesale out of SEO is arguing against the available data.
A measurement model that survives scrutiny
Baseline before you touch anything. Fix a prompt set that reflects real buying questions, sample it across the platforms your buyers use, and record the result. A baseline taken after work has started is not a baseline.
Instrument the surfaces separately. ChatGPT, Google AI Mode, Perplexity, Gemini and Copilot retrieve and cite differently. Reporting them as one number hides where you are winning and losing.
Hold a control where you can. Otterly.ai published an experiment adding the year to titles across 11 pages that showed apparent citation growth of 56 to 61%. One outlier drove 86 to 93% of the gain, and an untouched control group rose by a similar amount. The authors concluded the effect could not be separated from noise. That is the standard of honesty this field needs more of.
Separate leading from lagging. Citation share moves in weeks. Pipeline moves in quarters. Report both, and label which is which.
The cost side
Monitoring is now a real line item. Vendor-published pricing at the time of writing:
| Tool | Coverage | Public pricing |
|---|---|---|
| Otterly.ai | ChatGPT, AI Overviews, AI Mode, Perplexity, Copilot, Gemini | from $29/mo |
| Ahrefs Brand Radar | AI Overviews and AI Mode, ChatGPT, Copilot, Gemini, Perplexity, Grok | $398/mo select, $699/mo all platforms |
| Semrush AI Visibility | Prompt-level visibility, AI market share | Included in Semrush tiers |
| Profound, Peec AI, Evertune, Scrunch | Varying platform coverage and sampling depth | Sales-gated |
Pricing changes and we do not resell any of these. Worth stating plainly: no independent third-party accuracy audit of these tools exists that we could verify. Treat their numbers as directionally useful, not as ground truth.
What we will not claim
We do not guarantee citations, recommendations, or AI rankings. No one can. Retrieval is model-mediated, non-deterministic and changes without notice.
We do not present share of answer as revenue.
We do not attribute branded search lift to GEO without stating that the link is correlational.
What to ask before you buy
Ask any GEO vendor these five questions. The answers are diagnostic.
- What is my baseline, and when was it taken relative to the work starting?
- How many times do you sample each prompt, and across which models?
- Which of your reported numbers are causal and which are correlational?
- What did not work in the last engagement you ran?
- Are you asking me to block or allow AI crawlers, and can you explain the difference between the training decision and the retrieval decision?
A vendor who cannot answer the fifth question is the most expensive kind. Blocking GPTBot is a legitimate content-licensing choice. Blocking OAI-SearchBot removes you from ChatGPT search entirely. Confusing the two is the most common and costly error in the field, covered in detail in the GEO methodology.
Where this fits
ROI measurement is Phase 05 of our five-phase approach, and it feeds the next Phase 01. Read the methodology for how the phases depend on each other, the glossary for precise definitions of share of answer and citation share, and the FAQ for the questions clients ask most. If you want the tooling and reading list, see resources.