Five phases. One loop. Each phase produces the input the next phase consumes, and the last phase re-baselines the first.
This is the operational detail behind the summary on the GEO pillar page. It is written for two readers at once: an executive deciding whether we know what we are doing, and a practitioner who wants to run the work themselves. Both should find enough here to act on.
The order is not arbitrary, and it is not a marketing sequence. It is a dependency chain.
You cannot engineer answers for prompts you have not measured, so the audit comes first. You cannot build corroboration around an entity that machines cannot resolve, so the entity work comes before the citation work. You cannot judge whether any of it worked without a baseline that was fixed before you touched anything. Run the phases out of order and you will spend money on content that nobody retrieves, or earn mentions that get attached to the wrong company.
It is also a loop rather than a project. Models are retrained, retrieval systems change, competitors publish, and citations rot at a rate that surprises people. Phase 05 does not end the engagement. It produces the evidence that starts the next Phase 01.
How to read this page
Every phase chapter is built identically, so you can read across phases as well as down them.
| Subsection | What it contains |
|---|---|
| Goal | The one thing the phase exists to change |
| Entry criteria | What must be true before starting. If it is not true, you are in the wrong phase |
| Inputs | The specific artefacts and access the phase consumes |
| What we actually do | The procedure. Commands, file examples, schema, spreadsheet structures, prompt sets. This is the long part |
| Outputs & deliverables | The named artefacts a client receives and owns |
| Tools | What we use, why, and what it cannot tell you |
| Success metrics | What "done" looks like, numerically where numbers are honest |
| Worked example (illustrative) | One invented company carried across all five phases |
| Common mistakes | Four to six failure modes, each with the fix |
Three reading paths. Executives should read each Goal, Outputs and Success metrics, then the closing section on limits. Marketing leads should read Phases 01, 03 and 05 in full. Engineers and SEO teams should read Phases 01 and 02 in full and treat the rest as context.
Terms in bold-linked form on first use point at their definition in the GEO glossary.
The loop
flowchart LR
P1["01 AI Visibility Audit"] --> P2["02 Entity and Authority Foundation"]
P2 --> P3["03 Answer Engineering"]
P3 --> P4["04 Citation and Consensus Building"]
P4 --> P5["05 Monitor Measure and Compound"]
P5 -.->|"quarterly re-baseline"| P1
P5 -.->|"new prompt clusters"| P3
P5 -.->|"entity drift detected"| P2
P4 -.->|"corroboration gaps"| P2
P3 -.->|"unanswerable prompts"| P1
The five phases with their feedback edges. The dotted lines are the ones that make it a loop: measurement sends work backwards as often as it sends work forwards.
Read the dotted edges carefully, because they are where most of the year-two work actually comes from. Phase 05 finding that a competitor now owns a prompt cluster is a Phase 03 instruction. Phase 05 finding that a third-party directory has reverted your company description is a Phase 02 instruction. Phase 04 finding that no independent source will corroborate a claim is usually a sign the claim was never true enough to publish.
Typical durations, and why they vary
timeline
title Typical phase durations for a mid-size firm
section Weeks 1 to 4
Phase 01 AI Visibility Audit : Prompt set design : Baseline sampling : Access and rendering audit : Entity gap analysis
section Weeks 3 to 12
Phase 02 Entity and Authority Foundation : Disambiguation : Canonical description rollout : Structured data : Third-party reconciliation
section Weeks 6 to 20
Phase 03 Answer Engineering : Fan-out mapping : Answer-first rewrites : New cluster publication : Schema alignment
section Month 3 onward
Phase 04 Citation and Consensus Building : Entity distribution : Association and registry work : Corroboration outreach
section Month 2 onward and continuous
Phase 05 Monitor Measure and Compound : Monthly sampling : Link health : Quarterly re-baseline
Indicative ranges only. Phases overlap deliberately, and the tail of Phase 04 has no natural end.
Be sceptical of anyone who gives you a fixed calendar for this work, including us. Three things move the dates by months in either direction.
How broken the access layer is. A robots.txt fix can change what ChatGPT sees within days. A client-side-rendered site that needs an engineering release train to produce server-rendered text can take a quarter before Phase 03 has anywhere to publish.
How contested the entity is. A company with a unique name, one office and a clean Business Profile finishes Phase 02 in three weeks. A company that rebranded twice, operates under two legal entities and shares a name with a consumer brand can spend three months on disambiguation alone.
Who owns the third-party records. Association listings, registry entries and legacy directory profiles are frequently controlled by someone who left the company. Recovering access is unglamorous, slow, and not on anyone's critical path until it is.
Phases overlap. We start Phase 03 content design while Phase 02 reconciliation is still running, because reconciliation is mostly waiting on other people's ticket queues. What we do not do is start Phase 04 before Phase 02 is substantially complete. Earning mentions for an entity that is still ambiguous distributes the ambiguity.
Phase 01: AI Visibility Audit
Goal
Replace opinion with a baseline. At the end of Phase 01 the client knows, with a documented sampling method, how often they are named and linked in AI answers across the engines that matter to their buyers, how that compares with a named competitor set, which retrieval agents can and cannot reach their site, and which facts about their company the machine-readable web currently gets wrong. Almost nobody arrives with this. Most arrive with a screenshot of one ChatGPT answer that either flattered them or frightened them, which is not a measurement of anything.
Entry criteria
- A named decision-maker who can approve changes to robots.txt, DNS, CMS templates and third-party profiles. Without this the audit produces a report nobody can act on.
- Agreement on the buyer definition. GEO measurement is prompt-based, and prompts come from a specific buyer with a specific problem.
- A named competitor set of three to six companies. Not aspirational competitors. The ones that show up in the same shortlists.
- Read access to Google Search Console, Bing Webmaster Tools and server or CDN logs. Log access is the single most valuable input and the one most often refused on the first ask.
Inputs
- Primary domain and every alternate domain, including legacy domains from previous brand names.
- Full sitemap and, where available, a URL inventory with page templates identified.
- Current
robots.txt, CDN bot-management rules, and WAF configuration. - Twelve months of Search Console query and page data, plus Bing Webmaster Tools equivalents.
- Raw server or CDN access logs, minimum 30 days, ideally 90.
- A list of every third-party profile the company knows about: LinkedIn, Crunchbase, Google Business Profile, association directories, review platforms, registries.
- Sales-team language: the actual questions buyers ask on discovery calls, in their words.
What we actually do
1. Build the prompt set. This is the foundation of everything measured afterwards, and it is the step most often done badly. A prompt set is a fixed, versioned list of natural-language questions that a real buyer would type. It is not a keyword list, and converting a keyword list into questions produces a prompt set that measures nothing useful.
We build across five categories, and we track the mix deliberately because each one tells you something different.
| Category | What it measures | Typical share of set | Example shape |
|---|---|---|---|
| Unbranded discovery | Whether you exist in the consideration set at all | 30–35% | "who are the best X consultants in Y" |
| Comparison | Whether you survive a head-to-head | 20–25% | "X vs Y for Z", "alternatives to X" |
| Branded | What the engines say about you when asked directly | 15–20% | "is X any good", "what does X do" |
| Problem-led | Whether you are retrieved by symptom rather than category | 20–25% | "our submission was rejected for Z, what now" |
| Regulatory or risk | Whether you appear in the high-anxiety queries that drive urgent buying | 10–15% | "what are the penalties for Z in Ontario" |
A working set for a mid-size firm is 60 to 120 prompts. Below 40, engine variance swamps the signal. Above about 150, the sampling cost stops paying for itself and the marginal prompts are usually rephrasings.
Sources for prompts, in order of value: transcripts of sales discovery calls, the support inbox, Search Console queries with question shapes, community threads where your buyers argue, and last, competitor content titles. We do not generate the prompt set with an LLM. Models generate plausible-sounding prompts that no human has ever typed, and the resulting baseline measures a fictional buyer.
2. Decide the sampling protocol before running anything. Generated answers are non-deterministic. The same prompt, the same engine, the same day, produces different text and often different citations. This is not a flaw in your measurement; it is a property of the system. Any protocol that runs a prompt once and records the result is producing an anecdote.
Our default: each prompt, on each engine, five runs, in fresh sessions with personalization and memory disabled, geolocated to the client's primary market. Five is a compromise between cost and stability. Evertune's product samples each prompt 100 times across 11 models precisely because variance is large; we treat five as the working floor and raise it to ten for the prompts a client is making budget decisions on.
Record for each run: engine, model version if exposed, timestamp, geography, whether the brand was named, whether the brand was linked, the position of the mention in the answer, the full list of cited domains, and the verbatim sentence containing the mention. That last field is the one people skip and later wish they had, because framing matters. Being named as "a strong option for Class II submissions" and being named as "a smaller alternative to the national firms" are different commercial outcomes.
3. Calculate share of answer explicitly. Share of answer is the headline metric and it is frequently quoted without a definition, which makes it useless for comparison. Ours is stated in full, including the denominator.
Let P = number of prompts in the set
Let R = number of runs per prompt per engine
Let mention(p, r) = 1 if the brand is named in run r of prompt p, else 0
Let cited(p, r) = 1 if the brand's domain is linked in run r of prompt p, else 0
Σ(p=1..P) Σ(r=1..R) mention(p, r)
Share of answer (%) = ------------------------------------ × 100
P × R
Σ(p=1..P) Σ(r=1..R) cited(p, r)
Citation rate (%) = --------------------------------- × 100
P × R
number of prompts where Σ(r) mention(p, r) ≥ 1
Prompt coverage (%) = ------------------------------------------------ × 100
P
Three notes that stop this being misread. Share of answer is computed per engine and never averaged across engines into a single number, because the engines differ enough that the average describes nobody. Prompt coverage and share of answer diverge in a diagnostically useful way: high coverage with low share means you appear inconsistently and are probably a marginal candidate the model includes when it has room. Low coverage with high share means you own a narrow set of prompts and are absent everywhere else.
Where prompts differ in commercial value, we also report a weighted variant, with weights assigned by the client's sales team, not by us:
Weighted share of answer (%) = [ Σ(p) w(p) × ( Σ(r) mention(p, r) / R ) / Σ(p) w(p) ] × 100
4. Benchmark competitors on the identical set. Same prompts, same runs, same day. A competitor benchmark run a week later on a re-typed prompt list is not a benchmark. The working spreadsheet has one row per prompt and these columns:
| Column | Contents |
|---|---|
prompt_id |
Stable identifier, survives rewording |
prompt_text |
Verbatim |
category |
One of the five taxonomy buckets |
weight |
Commercial priority, 1 to 5, set by sales |
engine |
ChatGPT / AI Overviews / AI Mode / Gemini / Claude / Perplexity / Copilot |
run_index |
1 to R |
run_timestamp |
ISO 8601 |
geo |
Sampling location |
client_named |
0 / 1 |
client_linked |
0 / 1 |
client_sentence |
Verbatim mention text |
competitor_a_named … competitor_f_named |
0 / 1 per competitor |
cited_domains |
Pipe-delimited list of every domain cited in the answer |
answer_hash |
Hash of the response text, to detect duplicate captures |
From cited_domains you get the second most useful artefact of the whole audit: a frequency table of which domains the engines actually trust for your buyers' questions. That table sets the target list for Phase 04, and it is almost always surprising. Trade associations, one specialist blog nobody in the client's marketing team had heard of, and a government page tend to outrank the publications the client has been pitching for years.
5. Audit crawler access. Separate the training decision from the retrieval decision. Read robots.txt line by line against the documented agent table.
| Vendor | Agent | Purpose | Notes |
|---|---|---|---|
| OpenAI | GPTBot |
Model training | Opt out via robots.txt. IPs at openai.com/gptbot.json |
| OpenAI | OAI-SearchBot |
Surfaces sites in ChatGPT search | Blocking this removes you from ChatGPT search. IPs at openai.com/searchbot.json |
| OpenAI | ChatGPT-User |
User-triggered fetch | OpenAI describes it as non-automatic |
| OpenAI | OAI-AdsBot |
Ad landing page validation | Not used for training |
| Anthropic | ClaudeBot |
Training data | Supports non-standard Crawl-delay. IPs at claude.com/crawling/bots.json |
| Anthropic | Claude-User |
User-directed page fetch | |
| Anthropic | Claude-SearchBot |
Indexes content to improve search quality | |
| Perplexity | PerplexityBot |
Surfaces and links sites in Perplexity | Does not crawl for model training |
| Perplexity | Perplexity-User |
User-triggered fetch | Perplexity's own docs state it generally ignores robots.txt because the fetch is user-initiated |
Googlebot |
Search index | Gates AI Overviews and AI Mode eligibility | |
Google-Extended |
Opt-out token for AI training and grounding in some Google systems | Does not affect Search indexing | |
| Microsoft | bingbot |
Bing index |
Then verify behaviour rather than trusting the file, because CDN rules and WAFs override robots.txt and nobody remembers they exist. We fetch key templates with each documented user-agent string and record the status code:
for UA in \
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko); compatible; GPTBot/1.4; +https://openai.com/gptbot" \
"Mozilla/5.0 (Macintosh; Intel Mac OS X 10_15_7) AppleWebKit/537.36 (KHTML, like Gecko) Chrome/131.0.0.0 Safari/537.36; compatible; OAI-SearchBot/1.4; +https://openai.com/searchbot" \
"Mozilla/5.0 AppleWebKit/537.36 (KHTML, like Gecko; compatible; PerplexityBot/1.0; +https://perplexity.ai/perplexitybot)"
do
printf '%s\t' "$(curl -s -A "$UA" -o /dev/null -w '%{http_code} %{time_total}s' https://example.ca/services/)"
printf '%s\n' "${UA:0:40}"
done
A 403 or a challenge page returned to OAI-SearchBot while Googlebot gets a 200 is a bot-mitigation rule, not a robots.txt problem, and it is invisible to every SEO crawler on the market. We then confirm against logs, which tell you what is actually happening rather than what should be:
awk -F'"' '{print $6}' access.log \
| grep -Eo 'GPTBot|OAI-SearchBot|ChatGPT-User|ClaudeBot|Claude-SearchBot|Claude-User|PerplexityBot|Perplexity-User|Googlebot|bingbot' \
| sort | uniq -c | sort -rn
Zero hits from OAI-SearchBot over 90 days on a site that is not blocking it is its own finding, and usually means the site has never been discovered as a candidate at all.
6. Check rendering and text availability. Google's guidance is explicit that you should keep important content in text form. We test whether the answer to each priority prompt exists in the raw HTML, not just in the rendered DOM.
# Words present without JavaScript execution
curl -s https://example.ca/services/class-iii-licensing/ \
| sed -e 's/<script[^>]*>.*<\/script>//g' -e 's/<[^>]*>/ /g' \
| tr -s ' \n' ' ' | wc -w
Compare that count with the rendered word count from a headless browser. A gap above roughly 30% on a template means the substance is being injected client-side, and every retrieval system that does not execute JavaScript sees an empty shell. We also flag facts locked inside images or PDFs, infinite-scroll archives with no crawlable pagination, and nosnippet or max-snippet directives left over from a publisher-era policy, which restrict AI features exactly as they restrict snippets.
7. Run the entity gap analysis. For each of the following, record what the web currently says versus what is true: legal name, trading name, former names, founding year, headquarters city, office locations, employee count band, industry classification, services offered, named leadership, and the one-sentence description. Sources checked: the client's own site, LinkedIn, Crunchbase, Google Business Profile, the Knowledge Panel if one exists, Wikidata, Wikipedia, industry association directories, corporate registries, and the top ten third-party pages that mention the brand.
Then, separately, ask each engine a set of pure entity questions with no commercial framing: "what is X", "where is X based", "who founded X", "how many people work at X", "what does X specialize in". Record the answers verbatim. This is the fastest way to see the smear. When four engines give four founding years, you have found the Phase 02 work.
flowchart TD
A["Prompt set v1.0"] --> B["Sampling run, 5x per prompt per engine"]
B --> C["Share of answer and citation rate"]
B --> D["Cited domain frequency table"]
E["robots.txt and log audit"] --> F["Access findings"]
G["Render and text checks"] --> F
H["Entity questions to each engine"] --> I["Entity gap register"]
J["Third-party record inventory"] --> I
C --> K["Baseline report"]
D --> K
F --> K
I --> K
K --> L["Phase 02 scope"]
K --> M["Phase 03 prompt clusters"]
D --> N["Phase 04 target domains"]
Phase 01 produces four independent evidence streams that converge into one baseline and split again into three downstream scopes.
Outputs & deliverables
- Prompt Set v1.0 — versioned spreadsheet, five categories, weighted by sales priority. The client owns this file permanently.
- Baseline Visibility Report — share of answer, citation rate and prompt coverage per engine, with the sampling protocol stated on the same page as the numbers.
- Competitive Benchmark — the same three metrics for every competitor in the set, on the identical prompts and runs.
- Cited Domain Frequency Table — every domain the engines cited for your buyers' questions, ranked.
- Crawler Access Report — robots.txt line-by-line assessment, live user-agent fetch results, and 90-day log evidence.
- Rendering & Text Availability Findings — per template, with the raw-versus-rendered word gap.
- Entity Gap Register — every field where a third-party record disagrees with the truth, with the owner and the correction route for each.
- Prioritized Fix List — everything found, sorted by effort against expected effect, with the access-layer items at the top because they are cheap and fast.
Tools
We use Ahrefs Brand Radar or Otterly.ai for cross-engine sampling, Screaming Frog or Sitebulb for crawl and rendering comparison, Search Console and Bing Webmaster Tools for indexation truth, and raw log analysis for crawler behaviour. Where a client has budget for it, Evertune's high-sample-count approach is the most rigorous variance treatment we have seen commercially.
Two honest caveats. No independent third-party accuracy audit of any AI-visibility tool exists that we have been able to verify. Every vendor reports on its own coverage using its own sampling, and none of them publish an error rate. We use them because a systematic sample is better than no sample, not because their numbers are ground truth. Second, pricing moves and we do not resell any of them. Ahrefs Brand Radar publishes $398/month for selected platforms and $699/month for all; Otterly.ai starts at $29/month; Profound, Peec AI, Evertune and Scrunch are sales-gated. The fuller comparison is in the resource library.
Where tools disagree with our manual sampling, we report both and say which we trust more, and why.
Success metrics
Phase 01 is done when all of the following are true.
- Prompt set built, categorized, weighted, and signed off by someone in sales.
- Every prompt sampled at least five times on every in-scope engine, with the raw response text retained.
- Share of answer, citation rate and prompt coverage calculated per engine for the client and every competitor.
- 100% of AI-relevant user agents tested live, with status codes recorded, and cross-checked against a minimum 30 days of logs.
- Every priority template checked for raw-HTML text availability.
- Entity gap register complete for all twelve fields across all inventoried sources.
- Fix list prioritized and accepted.
Worked example (illustrative)
Halcourt Regulatory Partners is an invented company used throughout this page to keep the five phases concrete. It is not a client and its numbers are illustrative.
Halcourt is an 85-person regulatory affairs consultancy headquartered in Ottawa with offices in Toronto and Calgary. It helps medical device and diagnostics manufacturers obtain Health Canada licences and FDA clearance. It was founded in 2009 as Halcourt & Vance Consulting and rebranded to Halcourt Regulatory Partners in 2021. Its buyers are VPs of Regulatory Affairs and COOs at manufacturers with 50 to 500 staff.
The audit built 84 prompts. Examples, one per category:
- Unbranded discovery: "who can help a medical device manufacturer get a Class III licence in Canada"
- Comparison: "independent regulatory consultant versus building an in-house regulatory affairs team"
- Branded: "what does Halcourt Regulatory Partners do"
- Problem-led: "Health Canada issued a screening deficiency notice on our device licence application, what happens next"
- Regulatory or risk: "what are the Canadian medical device licence renewal requirements and penalties for lapse"
Illustrative baseline, five runs per prompt per engine:
| Engine | Share of answer | Citation rate | Prompt coverage |
|---|---|---|---|
| ChatGPT | 3.6% | 1.2% | 11% |
| Google AI Overviews | 6.9% | 4.5% | 19% |
| Gemini | 5.2% | 3.8% | 15% |
| Perplexity | 9.0% | 8.3% | 26% |
| Claude | 4.3% | 2.4% | 13% |
The strongest competitor in the set scored roughly four times higher on unbranded discovery and about the same on branded prompts. That pattern is diagnostic: Halcourt was not unknown, it was not being retrieved as a candidate.
Three access findings explained most of it. The CDN returned 403 to OAI-SearchBot under a bot-mitigation rule added during a scraping incident in 2024, while Googlebot was allowlisted. PerplexityBot was disallowed in robots.txt by a line added at the same time as GPTBot, with no evidence anyone had distinguished them. And the entire services section rendered its body copy client-side, leaving raw HTML word counts under 120 on pages that displayed 1,400 words.
The entity questions produced three different founding years across five engines, two of which described Halcourt as a law firm.
Common mistakes
Sampling each prompt once. The failure mode is a baseline made of noise, followed by a month-two report that shows dramatic improvement or collapse that never happened. The fix: fix R before you start, never change it mid-programme, and publish it next to every number.
Building the prompt set from keywords. Keyword-derived prompts measure the questions your SEO tool knows about, not the questions your buyers ask. Coverage looks respectable while the prompts that generate revenue go unmeasured. The fix: start from sales call transcripts and the support inbox. If a prompt cannot be traced to a human who asked something like it, cut it.
Trusting robots.txt instead of testing behaviour. Robots.txt says Allow, the CDN returns 403, and the audit reports a clean access layer. The fix: live user-agent fetches on every priority template, cross-checked against logs. Behaviour beats configuration.
Auditing only ChatGPT. The engines diverge widely on the same prompt set. A single-engine audit produces a strategy tuned to one retrieval stack. The fix: sample every surface where your buyers actually are, and report per engine, never averaged.
Skipping the competitor run to save budget. Share of answer without a comparison is a number with no scale. A 6% share is excellent in a fragmented market and a crisis in a concentrated one. The fix: the competitor run is not optional. It is the axis.
Letting the audit finish before securing change authority. The report lands, the recommendations require an engineering release and a CDN rule change, and there is no owner. The fix: the entry criteria exist for a reason. Get the named approver before the first fetch.
Phase 02: Entity & Authority Foundation
Goal
Make the company a resolvable entity. At the end of Phase 02 a retrieval system can determine who you are, what you do, where you operate and that you are not somebody else with a similar name, and every significant machine-readable record on the web agrees. This is the phase clients find least exciting and the one that changes outcomes most, because everything downstream attaches to the entity. Corroboration earned for an ambiguous entity is corroboration distributed across several partial identities, none of which is you.
The canonical name of this phase is Entity & Authority Foundation.
Entry criteria
- Phase 01 entity gap register complete.
- A decision-maker who can settle naming questions, including the legal-suffix rule and the fate of legacy brand names. This is frequently a founder conversation, not a marketing one.
- Access, or a recovery path, to every third-party profile in the inventory.
- Developer capacity to deploy structured data and template changes.
Inputs
- Entity gap register from Phase 01.
- Corporate registry filings, legal entity names, and any operating-name registrations.
- Complete inventory of third-party profiles with credentials or account-recovery routes.
- Leadership and subject-matter-expert list with credentials, licence numbers, publication records and professional profiles.
- Brand guidelines, including how the name is written in running text.
- Existing structured data, extracted from the live site rather than from the CMS configuration.
What we actually do
1. Resolve the entity collisions. Start with an inventory of everything a machine might confuse you with. Search the exact brand name, the brand name without its suffix, the founder surname, and the legacy names, across web search, LinkedIn company search, Crunchbase, corporate registries and Wikidata. Record every collision and classify it: same-name different-industry, same-name same-industry, former-self, or subsidiary confusion.
Same-name different-industry collisions are usually survivable and are handled by making the industry context unmissable in every description. Same-name same-industry collisions require a naming decision, and sometimes the right answer is that the client needs a modifier in their trading name. Former-self collisions are the most common and the most fixable: the old name persists across directories, PDFs, press archives and the client's own legacy URLs.
2. Fix the canonical name, once. Write the rule down and apply it without exception:
- One canonical form in running text, including or excluding the legal suffix. Pick one. Never both.
- One short form, defined explicitly, used consistently after first mention.
- Legacy names appear only in a single "formerly known as" statement on the About page, and in nowhere else that we control.
- The rule covers the
<title>pattern,Organization.name,Organization.legalName,Organization.alternateName, every profile display name, and email signatures.
3. Write the canonical description. One sentence. It answers what the company does, for whom, and where. It contains no adjectives that a competitor could also claim, and no words a buyer would not use.
Then the discipline that makes it work: that sentence is used verbatim, character for character, everywhere. Website meta description, homepage sub-headline, Organization.description in JSON-LD, LinkedIn About, Crunchbase overview, Google Business Profile description, every association directory listing, the boilerplate at the bottom of every press release, and the speaker bio the CEO sends to conferences.
This looks like a trivial copywriting task. It is the highest-yield hour in most engagements, and the reason is mechanical. Generative systems resolve conflicting claims by hedging, by picking the most authoritative source, or by omitting the entity entirely. Eleven sources saying the same sentence produces a fact. Eleven sources saying eleven things produces nothing retrievable, and the confident competitor is named instead.
4. Ship Organization structured data. Google's documentation is clear that there is no special schema required for AI features, and we say so on the pillar page. Structured data is not a magic token. What it does is state your identity in an unambiguous machine-readable form, and sameAs is the property that binds your identity to the profiles that corroborate it.
{
"@context": "https://schema.org",
"@type": "Organization",
"@id": "https://halcourt.example.ca/#organization",
"name": "Halcourt Regulatory Partners",
"legalName": "Halcourt Regulatory Partners Inc.",
"alternateName": "Halcourt",
"url": "https://halcourt.example.ca",
"description": "Halcourt Regulatory Partners is an Ottawa-based regulatory affairs consultancy that helps medical device and diagnostics manufacturers obtain Health Canada licences and FDA clearance.",
"foundingDate": "2009",
"numberOfEmployees": { "@type": "QuantitativeValue", "value": 85 },
"areaServed": ["CA", "US"],
"knowsAbout": [
"Medical device regulation",
"Health Canada medical device licensing",
"FDA 510(k) clearance",
"ISO 13485 quality management systems"
],
"address": {
"@type": "PostalAddress",
"streetAddress": "000 Example Street, Suite 000",
"addressLocality": "Ottawa",
"addressRegion": "ON",
"postalCode": "K1P 0A0",
"addressCountry": "CA"
},
"sameAs": [
"https://www.linkedin.com/company/example-halcourt/",
"https://www.crunchbase.com/organization/example-halcourt",
"https://www.example-association.ca/directory/halcourt",
"https://www.youtube.com/@example-halcourt",
"https://x.com/example_halcourt"
]
}
Rules we hold to. Every sameAs URL must be a profile the client verifiably controls or is verifiably listed on, and must resolve with a 200. A dead sameAs is a contradiction, not a signal. The description is the canonical sentence, unmodified. knowsAbout uses the vocabulary buyers and regulators use, not internal service names. And the whole block matches visible text on the page, which is one of the few structured-data instructions Google states directly for AI features.
5. Reconcile third-party records. Build a reconciliation sheet and work it like a ticket queue. One row per record, one column per field:
| Column | Contents |
|---|---|
source |
LinkedIn / Crunchbase / Business Profile / association / registry / review site |
url |
Direct link to the record |
field |
name / description / founding year / HQ / employees / industry / URL |
current_value |
Verbatim, as it appears today |
canonical_value |
What it must say |
status |
correct / wrong / missing / disputed |
access |
have credentials / recoverable / requires vendor support / no route |
owner |
Named person responsible for the change |
submitted |
Date the correction was submitted |
verified |
Date the change was observed live |
recheck |
Next audit date |
The verified column matters because a meaningful number of directory edits silently revert, and a correction you submitted but never confirmed is not a correction. The recheck column is what turns this from a project into a maintained asset, and it is one of the inputs to the Phase 05 cadence.
Priority order, based on how heavily these sources appear in cited-domain tables: Google Business Profile, LinkedIn, the industry associations your buyers recognize, corporate registries, Crunchbase, then the long tail of directories. We do not chase directory listings for their own sake. A hundred low-quality listings is a hundred future contradictions to maintain.
6. Decide about Wikidata honestly. Wikidata is a structured, openly licensed knowledge base and it is genuinely useful when an entity belongs there. It is appropriate when the organization already has independent coverage that establishes it as a documented thing: press coverage, registry entries, association membership, published research, notable products.
It is not appropriate as a manufactured shortcut. Creating an item for a company with no independent sourcing invites deletion, and a deleted item is worse than no item because the deletion discussion becomes part of the public record. The same applies with greater force to Wikipedia, where notability standards are enforced by volunteers who are extremely good at spotting agency-written articles. We do not manufacture notability. If the sourcing is not there, the correct action is Phase 04 work first, and Wikidata later, if ever.
Where an item is warranted, we make sure the statements carry references, that the item links to the official website, and that the label and description match the canonical forms.
7. Do the author entity work. For firms whose credibility rests on named experts, the people are entities too, and an author with no verifiable footprint contributes nothing to trust.
For each named expert: a real author page at a stable URL, Person structured data with sameAs pointing to LinkedIn, professional-body listings, ORCID where relevant, and licence or registration records; visible credentials with issuing body and, where public, a registration number; a list of published work; and author on every article they wrote, pointing at the same @id.
{
"@context": "https://schema.org",
"@type": "Person",
"@id": "https://halcourt.example.ca/team/example-author/#person",
"name": "Example Author",
"jobTitle": "Principal, Medical Device Regulatory Affairs",
"worksFor": { "@id": "https://halcourt.example.ca/#organization" },
"alumniOf": "Example University",
"knowsAbout": ["Health Canada device licensing", "ISO 13485", "FDA 510(k)"],
"sameAs": [
"https://www.linkedin.com/in/example-author/",
"https://orcid.org/0000-0000-0000-0000"
]
}
8. Verify, do not assume. Re-fetch the live pages and validate the emitted JSON-LD from rendered output. CMS plugins frequently emit a second, conflicting Organization block, and two organizations on one page is a disambiguation failure you created yourself.
flowchart TD
A["Collision inventory"] --> B["Canonical name rule"]
B --> C["Canonical description, one sentence"]
C --> D["Organization JSON-LD with sameAs"]
C --> E["Third-party reconciliation queue"]
D --> F["Single resolvable entity"]
E --> F
G["Author pages and Person schema"] --> F
H{"Independent sourcing exists?"} -->|Yes| I["Wikidata item with references"]
H -->|No| J["No item. Return after Phase 04"]
I --> F
F --> K["Phase 03 can attach content to a known entity"]
F --> L["Phase 04 can attach citations to a known entity"]
Everything in Phase 02 converges on one question: can a machine resolve you to a single thing? Wikidata is a branch, not a step.
Outputs & deliverables
- Entity Definition Document — canonical name, short form, legacy-name policy, canonical description, approved boilerplate, and the rules for applying them.
- Structured Data Package — deployed
OrganizationandPersonJSON-LD, with a validation report from rendered output. sameAsRegister — every profile URL, verified live, with an owner.- Third-Party Reconciliation Sheet — every record, field, current value, canonical value, status, owner and recheck date.
- Author Entity Pack — author pages, credentials,
Personschema, and a byline policy. - Wikidata Assessment — a documented yes or no with the reasoning, and the sourcing gap if the answer is no.
- Entity Change Log — dated record of every correction submitted and verified, which becomes the evidence base for attribution discussions later.
Tools
Schema validation with Google's Rich Results Test and the Schema.org validator, understanding that neither validates truth. Rendered-DOM extraction to catch duplicate blocks. Google Business Profile, LinkedIn and Crunchbase admin interfaces. Corporate registry search for legal names. Wikidata's own query service to check for existing items and collisions.
The honest limit: none of these tools tell you whether an engine has actually updated its representation of you. There is no API that returns "here is what the model believes about your company." You infer it by re-running the Phase 01 entity questions and reading the answers. That inference is soft, and we present it as inference.
Success metrics
- 100% of controlled properties carrying the canonical description verbatim.
- 100% of
sameAsURLs resolving 200 and pointing to records that carry the canonical description. - Zero duplicate or conflicting
Organizationblocks in rendered output. - Zero fields in the reconciliation sheet with status
wrongand no owner. - 100% of named authors with an author page, credentials and
Personschema. - Entity questions re-run: the target is agreement on founding year, headquarters and industry classification across all sampled engines. This one is a target, not a promise, because the update horizon is outside anyone's control.
Worked example (illustrative)
Halcourt's collision inventory found three: a UK property firm trading as Halcourt, a US immigration attorney named Halcourt, and Halcourt's own former identity, Halcourt & Vance Consulting, which was still the name in two association directories and on 40-odd archived press pages.
The naming decision: "Halcourt Regulatory Partners" in all running text, "Halcourt" as the short form after first mention, "Halcourt Regulatory Partners Inc." reserved for legalName and contracts. The 2009–2021 name appears in exactly one sentence on the About page.
The canonical description, agreed in a 40-minute meeting after three weeks of email:
Halcourt Regulatory Partners is an Ottawa-based regulatory affairs consultancy that helps medical device and diagnostics manufacturers obtain Health Canada licences and FDA clearance.
Thirty-one words. It names the city, the discipline, the buyer and the two outcomes. It survives being lifted into an answer intact, which is the actual test.
Reconciliation found 22 records across 14 sources. Nine were wrong on at least one field, four still used the legacy name, and two described the firm as a law firm, which explained the two engines that had said the same thing. Six required account recovery, one of which took nine weeks because the registered contact had left in 2019.
Wikidata assessment: no item created. Halcourt had trade-press mentions but nothing that would survive a notability discussion. The assessment documented what sourcing would be needed and deferred the question to after Phase 04.
Author work covered six consultants. Two had no LinkedIn presence at all, which was raised as a business decision rather than a marketing one.
Common mistakes
Treating the canonical description as copywriting. Marketing rewrites it for each channel because repetition feels lazy. The contradiction is reintroduced by the people who fixed it. The fix: write the rule into the brand guidelines and make verbatim reuse an approval condition, not a suggestion.
Adding sameAs entries for profiles you do not control. A sameAs pointing at a stale profile with the old description actively asserts a contradiction. The fix: every sameAs URL must resolve 200 and carry the canonical description, or it comes out.
Manufacturing a Wikidata or Wikipedia presence. Deletion is the likely outcome, the discussion is public and permanent, and the reputational cost outlasts the item. The fix: earn the sourcing in Phase 04 first. Wikidata is a lagging indicator of notability, not a way to create it.
Shipping schema that contradicts visible text. The JSON-LD says 85 employees, the About page says "over 100," and the structured data is now a reason to distrust the page. The fix: schema is a machine-readable restatement of visible content. Diff them before deploying.
Letting the CMS emit a second Organization block. A theme and a plugin each emit one, they disagree, and the page describes two companies. The fix: validate rendered output, not source templates, after every plugin update.
Ignoring the people. The firm is treated as an entity while its named experts remain anonymous bylines, which removes the strongest credibility signal a professional-services firm has. The fix: author entity work is in scope, and its cost is mostly the experts' time, not budget.
Phase 03: Answer Engineering
Goal
Make sure that for every prompt that matters, an answer exists on your site, in text, near the top of a page, phrased so it can be lifted intact and attributed. Phase 03 converts the prompt set from a measurement instrument into a content specification. This is where content strategy stops being about terms and starts being about questions and their decompositions.
Entry criteria
- Phase 02 substantially complete. Publishing at volume against an unresolved entity spreads the ambiguity across more pages.
- Access-layer fixes from Phase 01 deployed. There is no point publishing for crawlers that get a 403.
- Prompt set v1.0 with sales weightings.
- Subject-matter experts committed to a review cadence. Answer engineering without expert input produces fluent content with nothing in it, which is the exact failure mode of the AI-content experiments that got deindexed.
Inputs
- Prompt set with categories and weights.
- Cited-domain frequency table from Phase 01, which shows what a winning answer currently looks like for these prompts.
- Full content inventory with current rankings, indexation status and last-modified dates.
- Entity Definition Document from Phase 02.
- Expert availability calendar and any proprietary data the firm holds.
What we actually do
1. Decompose the fan-out for every priority prompt. Google documents that AI Overviews and AI Mode may use a query fan-out technique, "issuing multiple related searches across subtopics and data sources," and describes AI Mode as making a plan, running searches and adjusting the plan based on what it finds. OpenAI documents that ChatGPT search "rewrites your query into one or more targeted queries."
Both statements mean the same operational thing. The engine is not searching for your buyer's question. It is searching for the four to eight sub-questions it derived from it. Optimizing for the literal prompt is the wrong target. See query fan-out for the definition.
We decompose manually and verify empirically. Take a priority prompt:
"We're a 60-person diagnostics manufacturer in Ontario. Our Class III device licence application got a screening deficiency notice. Who can help and how long will it take to fix?"
Decomposed, this is at least six distinct information needs:
- What is a Health Canada screening deficiency notice?
- What are the response deadlines and what happens if you miss them?
- What are the common causes of screening deficiencies for Class III devices?
- Who provides regulatory remediation support in Ontario or Canada?
- How much does that engagement typically cost and how long does it take?
- Should this be handled in-house or externally?
Six needs, and most firms have content for exactly one of them: number four, their service page. The service page cannot answer needs one, two, three, five or six, so it is never the source the engine reaches for on those sub-queries, and it therefore never enters the synthesis at all.
The empirical verification step matters. Run the parent prompt on an engine that exposes its searches, note the sub-queries it actually issues, and compare them with your manual decomposition. You will be wrong about two of six, and the two you were wrong about are usually the commercially interesting ones.
We record fan-out in a map with one row per sub-question and these columns: parent_prompt_id, sub_question, information_type (definition / procedure / criteria / comparison / cost / risk), existing_url, gap (none / partial / missing), target_url, owner, expert_reviewer, status.
2. Design the content architecture around clusters, not pages. A cluster is one parent page that owns the whole topic plus a set of child pages that each own exactly one sub-question. The parent links to every child with descriptive anchor text; every child links back and sideways to its two most-related siblings.
Semantic clustering means grouping by the concept a buyer holds in their head, not by your service taxonomy. Buyers do not think in service lines. They think "my application got rejected." A cluster organized around that state of the world outperforms one organized around "Regulatory Remediation Services," because the former matches the sub-questions and the latter matches your org chart.
Each child page has one job: answer one sub-question completely, in text, in the first 60 words, and then earn its length with detail, data and named expertise. If a page needs two H1-worthy answers, it is two pages.
3. Rewrite for answer-first extraction. This is a specific, teachable structure. The first paragraph after the H1 is a self-contained answer to the page's question. It is written so that if it were lifted out of the page and pasted into an AI answer with no other context, it would be correct, attributable and complete.
Before, which is how most professional-services pages are written:
In today's evolving regulatory environment, medical device manufacturers face increasing scrutiny. Our team brings decades of combined experience helping innovative companies navigate complex requirements. We understand that every submission is unique, and our tailored approach ensures that your regulatory strategy aligns with your commercial objectives. In this article, we will explore some of the considerations around screening deficiency notices.
That paragraph contains no facts. It cannot be extracted, because there is nothing in it to extract. It also delays the actual answer past the point where an extraction system stops reading.
After:
A screening deficiency notice is Health Canada's formal notification that a medical device licence application is incomplete and cannot proceed to review. The applicant must respond within the period stated in the notice. If the response is not filed in time, the application is cancelled and the manufacturer must submit and pay for a new application. The most common causes are incomplete quality management system evidence, missing labelling in both official languages, and clinical evidence that does not match the stated device classification. Verify current timelines and requirements against the Health Canada guidance in force at the time of your application.
Ninety-four words. Definition first, consequence second, causes third, verification caveat fourth. Every sentence stands alone. It is quotable in isolation, which is the entire objective.
4. Apply the GEO paper's tactic findings, carefully. The GEO paper (arXiv:2311.09735, KDD 2024) tested nine content modifications against its GEO-bench benchmark. The abstract states the method "can boost visibility by up to 40% in generative engine responses," and also that "the efficacy of these strategies varies across domains, underscoring the need for domain-specific optimization methods." The study predates AI Mode and the current model generation. Treat it as directional evidence about how synthesis systems select material, not as a 2026 ranking-factor list.
| Finding | Practical instruction |
|---|---|
| Quotations from relevant or expert sources — largest gain, around 40% | Every substantive page carries at least one attributed quotation, ideally from a named internal expert with a credential, or from the regulator's own text |
| Statistics and data points — around 30% | Replace qualitative claims with figures and sources. "Applications are frequently rejected" becomes a cited number or it comes out |
| Fluency and readability — around 28–30% | Short sentences, one idea per sentence, no clause stacking. Copy-edit as an optimization step, not a courtesy |
| Citing sources — around 27–28% | Link the primary source inline. Regulator pages, standards bodies, published research |
| Technical terminology — around 18–20% | Use the buyer's and the regulator's actual vocabulary, including document names and section numbers |
| Simplified language — around 14–15% | Explain the term on first use, then use it properly |
| Authoritative tone — around 10–13% | Real, but small. Do not substitute tone for substance |
| Unique word count — negligible | Do not chase vocabulary variety |
| Keyword stuffing — negative | Term-frequency optimization is a liability here, not a neutral habit |
The through-line: the tactics that helped are the ones that make a passage worth quoting. That is not a trick, and it is why we do not treat this table as a checklist to be gamed. A page with a fabricated statistic satisfies the letter of row two and fails at everything that matters.
5. Ship schema that matches visible text. Google's guidance for AI features says to ensure structured data matches visible text. That instruction is the whole discipline. Article with a real author reference into the Phase 02 Person graph, datePublished and dateModified that are true, and about terms that match the page's actual subject.
And a statement that most GEO guides still get wrong. FAQPage markup no longer produces a Google rich result. Google's documentation says, verbatim:
"As of May 7, 2026, FAQ rich results are no longer appearing in Google Search. We will be dropping the FAQ search appearance, rich result report, and support in the Rich results test in June 2026. To allow time for adjusting your API calls, support for the FAQ rich result in the Search Console API will be removed in August 2026."
HowTo rich results were removed earlier, in September 2023. FAQPage markup remains legitimate machine-readable semantics for other consumers, and shipping it is cheap and harmless. What it is not is a Google visibility tactic. Anyone selling FAQ schema as an AI-visibility service in 2026 is working from a 2023 playbook. We ship it where it describes real on-page Q&A content, and we tell clients plainly what it does and does not buy.
6. Instrument before publishing. Every new page gets registered in the prompt map against the sub-questions it is meant to answer, so Phase 05 can attribute movement to specific publications rather than to the programme in general. Pages published outside the map are not measurable and mostly should not exist.
7. Fix the old before making the new. Roughly half of Phase 03 output in a typical engagement is rewriting existing pages that already have authority and bury their answer in paragraph fourteen. Rewriting an indexed page with existing links is faster, cheaper and more reliable than publishing a new one, and it is consistently the higher-return half of the work.
flowchart TD
A["Priority prompt"] --> B["Manual fan-out decomposition"]
A --> C["Observed engine sub-queries"]
B --> D["Reconciled sub-question list"]
C --> D
D --> E{"Existing page answers it?"}
E -->|"Yes, but buried"| F["Answer-first rewrite"]
E -->|"Partially"| G["Extend and restructure"]
E -->|"No"| H["New child page in cluster"]
F --> I["Expert review"]
G --> I
H --> I
I --> J["Schema aligned to visible text"]
J --> K["Register in prompt map"]
K --> L["Publish and hand to Phase 05"]
Answer engineering starts from the decomposition, not from the page. Most output is repair, not new publication.
Outputs & deliverables
- Fan-Out Map — every priority prompt decomposed into sub-questions, each mapped to an owning URL.
- Content Gap Register — sub-questions with no adequate answer, prioritized by prompt weight.
- Cluster Architecture Diagram — parent and child pages with the internal link model.
- Answer-First Style Guide — the extraction structure, with before-and-after pairs drawn from the client's own pages.
- Rewritten Priority Pages — existing URLs restructured, with the pre-change version archived.
- New Cluster Pages — published, expert-reviewed, schema-aligned.
- Schema Implementation Notes — including the written FAQPage position, so nobody re-litigates it in six months.
- Prompt-to-URL Registry — the mapping Phase 05 measures against.
Tools
Search Console for existing query and page performance. A crawler for internal link structure and orphan detection. Engines that expose their search steps for fan-out verification. Readability scoring as a rough proxy for the fluency finding, treated as a signal rather than a target. A schema validator run against rendered output.
The limit worth stating: no tool tells you which sub-questions an engine will generate for a given prompt in general. You observe a sample of actual behaviour and generalize. Fan-out is not fully documented by any vendor beyond the descriptions quoted above, and anyone presenting a definitive fan-out model is inferring.
Success metrics
- 100% of priority prompts decomposed and reconciled against observed engine behaviour.
- 90% or more of sub-questions with a named owning URL.
- 100% of priority pages answering their question within the first 60 words.
- Every substantive page carrying at least one attributed quotation and at least one sourced statistic.
- Zero pages with schema contradicting visible text.
- Keyword-density optimization removed from the content brief template entirely.
- Directionally, over the following two to three months: increased prompt coverage on the clusters shipped, measured against the Phase 01 baseline. Directional, because publication does not control retrieval.
Worked example (illustrative)
Halcourt's 84 prompts decomposed into 341 sub-questions. Deduplicated, 206 were distinct. Of those, 38 had an adequate existing answer, 51 were partially covered, and 117 had nothing.
The single highest-weighted parent prompt was the screening-deficiency scenario above. Halcourt had one page for it: "Regulatory Remediation Services," 620 words, opening with "Halcourt's experienced team supports manufacturers through every stage." It answered none of the six sub-questions.
The cluster shipped as one parent, "Health Canada medical device licence deficiencies," and seven children, each owning one sub-question. The definitional child page was written by a named principal and opened with the 94-word passage quoted earlier. Two proprietary assets were built: an anonymized breakdown of deficiency causes across 140 submissions Halcourt had handled, and a plain-language decision guide on in-house versus external remediation.
The proprietary breakdown mattered more than anything else in the cluster. It was the only page in the set that contained information not already available on a regulator's website, which is the practical meaning of the paper's statistics finding.
Rewrites covered 22 existing pages. Halcourt's team initially resisted, on the grounds that the pages were already ranking. They were. They were also invisible in AI answers, because the answer was in paragraph nine.
FAQPage decision, documented: shipped where genuine Q&A existed on the page, with a note in the CMS explaining that it no longer produces a Google rich result and is not the reason for shipping it.
Common mistakes
Building one page per prompt. The engine searches sub-questions, so a page built for the parent prompt competes for nothing the engine actually issued. The fix: build for the decomposition. One page per sub-question.
Answer-first as a formatting exercise. A summary box is added at the top of an unchanged article, containing the same empty positioning language. The fix: the opening passage must be extractable and factually complete on its own. Test it by reading it in isolation and asking whether it would be a correct answer.
Chasing the paper's tactics mechanically. A statistic is inserted because statistics scored 30%, and it is made up or irrelevant. This is worse than nothing, because it makes the page a liability. The fix: the tactics describe what makes a passage genuinely quotable. If you do not have the data, get the data or write a different sentence.
Publishing at AI-assisted volume. Otterly.ai published roughly 1,000 AI-generated posts each on two fresh domains; both were algorithmically deindexed by Google with no manual action, and one fell from 1,629 to 15 daily impressions. The sample is two sites, so do not over-generalize. It does falsify the claim that scaled AI publishing is risk-free. The fix: expert-reviewed depth over volume. Every page has a named author who can defend it.
Retaining keyword-density targets in the brief. The GEO paper found keyword stuffing had a negative effect. Content teams keep the old template because it is what the tool outputs. The fix: delete the field from the brief.
Never rewriting anything. All budget goes to new pages while indexed pages with existing authority keep burying the answer. The fix: audit the existing corpus first. Repair usually beats publication.
Phase 04: Citation & Consensus Building
Goal
Make the claims on your own site true according to sources you do not control. Phase 04 builds the corroboration layer: independent references, consistent entity data on third-party platforms, and presence in the venues the engines actually cite for your buyers' questions. Owned content is necessary and, per Profound's data, is where the majority of citations land. It is not sufficient. A claim only you make is a claim. A claim eleven independent sources make is a fact.
Entry criteria
- Phase 02 complete enough that a new mention will attach to a resolved entity.
- Phase 03 has published something worth citing. Outreach without an asset is a request for a favour.
- Cited-domain frequency table from Phase 01, which is the target list.
- Executive time committed. Corroboration work runs on the credibility of named humans, and it cannot be fully delegated to an agency.
Inputs
- Cited-domain frequency table, ranked.
- Association and professional-body memberships, current and lapsed.
- Existing media relationships, speaking history, award entries.
- Proprietary data or research from Phase 03 that a journalist or association would find genuinely useful.
- Entity Definition Document, because every new mention should carry the canonical description.
- Customer and partner relationships that could support a case study or a joint publication.
What we actually do
1. Work the three tiers deliberately. The pillar page sets out the citation architecture: owned, verified, earned. Phase 04 is where the second and third tiers get built, and the mistake is over-investing in whichever is easiest.
| Tier | What it is | Control | Speed | What it does for you |
|---|---|---|---|---|
| Owned | Your site, documentation, knowledge hub | Total | Fast | Supplies the substance the engines quote |
| Verified | LinkedIn, Crunchbase, Business Profile, Wikidata, association registries | High | Medium | Confirms the entity exists and is what it says |
| Earned | Trade press, comparison pages, community discussion, analyst mentions | Low | Slow | Supplies the independence that makes the other two credible |
Profound's analysis of 11.84 billion citations across 3.02 million domains, April to July 2026, found roughly 57% of AI citations globally point to brand-owned domains, ranging from about 47% for ChatGPT to 69% for Gemini. Single-vendor, large-sample, unaudited, so hold the exact figures loosely. The directional reading is the useful part: owned content is the primary citation surface, and it is only trusted because the other two tiers corroborate the entity behind it.
2. Understand what meaningful velocity is. Velocity is the rate at which new independent references appear. It matters because a pattern of steady accumulation reads differently from a burst.
Meaningful velocity has four properties. It is steady rather than spiked. It is varied in source type: trade press, associations, community threads, partner content, conference listings. It is varied in anchor and phrasing, because thirty identical anchor texts in a month is a signature, not a pattern. And it is tied to real events: a publication, a hire, a data release, a speaking engagement.
Spikes look manufactured because they usually are. Forty new mentions in a fortnight, all from similar low-quality domains, all with near-identical phrasing, all with no corresponding real-world event, is the profile of a purchased campaign, and it is exactly the profile that both search and generative systems are built to discount. We have no documented evidence of how any specific engine treats such a pattern, and we do not pretend otherwise. What we can say is that the pattern carries risk for no reliable benefit, and that a company with a real corroboration footprint does not need it.
A workable target for a mid-size firm is a handful of genuine new references a month, sustained. Sustained is the operative word. Twelve months of three per month builds something a burst of forty does not.
3. Run the entity data distribution checklist. This is the verified tier, and it is a checklist because it is finite and completable.
- Google Business Profile: complete, categorized correctly, canonical description, every location, current hours, verified.
- Bing Places: the same, and almost universally neglected.
- LinkedIn company page: canonical description, correct industry, correct size band, correct HQ, website URL, employees actually associated with the page.
- Crunchbase: canonical description, founding year, HQ, category, leadership.
- Industry associations and professional bodies: every membership carrying a live directory listing with the canonical description and a working link.
- Corporate registries: legal name and address current.
- Standards and certification bodies: certification listings, where they publish public registries.
- Review platforms relevant to the sector, claimed and accurate.
- Conference and event profiles: speaker bios using the canonical boilerplate.
- Partner and vendor directories: partner listings on the sites of software or service partners.
- Chambers of commerce and regional economic development listings, which are frequently well-indexed and almost always out of date.
- Award and ranking bodies: any listing where the firm has genuinely been recognized.
Each row gets an owner and a recheck date, and feeds the Phase 05 cadence. This checklist alone typically moves branded-prompt accuracy more than any other single activity, because these are precisely the pages that engines reach for when asked "what is X."
4. Do digital PR as corroboration, not as link acquisition. The link-building framing produces the wrong behaviour: chase domain authority, place a paragraph with an anchor, count the link. For generative retrieval the useful question is different. Does this source, which the engines already cite for my buyers' questions, now contain an accurate statement about my company?
That reframing changes targeting. From the cited-domain table you know exactly which domains the engines pull from. A mention on the specialist association site that appears in 30% of your prompt sample's citations is worth more than a mention on a general business title with a higher domain rating and zero appearances.
It changes the pitch, too. The asset is the proprietary data from Phase 03, offered as something a publication can report on, or a named expert offered as a source with verifiable credentials. It changes what "success" means: a mention with no link that describes the company accurately in an already-cited source is a genuine win, and a link from a domain the engines never cite is close to worthless here.
Practically: comparison and "best X" pages on third-party domains are disproportionately important, because they dominate commercial prompts. Getting accurately included in an independently maintained comparison, including correcting an existing listing that has you wrong, is often the highest-value single action in Phase 04. Community venues matter too, on the same terms: participate as a named expert answering questions, never as a covert promoter. The failure mode there is severe and permanent.
5. Be careful about paid placement. Some of what is sold as GEO citation-building is paid placement in listicles and comparison pages. State the position clearly.
Organic citations in AI answers are not purchasable inventory. There is no mechanism to buy inclusion in ChatGPT's organic recommendations, and a vendor who claims otherwise should be asked to document it. Advertising products inside AI surfaces exist, are labelled as advertising, and are a legitimate media buy that has nothing to do with organic citation.
Paid inclusion in a third-party listicle occupies the grey area. Sometimes it is a transparent sponsorship on a page that also carries editorial listings. Sometimes it is a page that exists only to sell placements, has no readership, and confers nothing. Two tests we apply: would this page exist if nobody paid for placement, and does it appear in the cited-domain table? If the answer to both is no, it is not worth buying at any price. And undisclosed paid placement is a disclosure problem before it is a marketing one.
6. Close the loop back into Phase 02. Every new earned mention is checked against the canonical description. Publications paraphrase, and paraphrases drift. Where a significant source has the description materially wrong, we request a correction. Where it is a minor variation, we log it and leave it. Perfect uniformity is neither achievable nor necessary; what matters is that no widely-cited source asserts something false about the entity.
flowchart LR
A["Cited domain frequency table"] --> B["Target source list"]
B --> C["Verified tier checklist"]
B --> D["Earned tier outreach"]
E["Proprietary data from Phase 03"] --> D
F["Named experts with credentials"] --> D
C --> G["Consistent entity data"]
D --> H["Independent accurate mentions"]
G --> I(("Corroboration"))
H --> I
I --> J["Engines can verify claims off-domain"]
H -.->|"description drift found"| K["Phase 02 correction"]
Phase 04 targets the domains the engines already trust, and routes any description drift straight back into Phase 02.
Outputs & deliverables
- Corroboration Target List — ranked by observed citation frequency for the client's prompt set, not by domain authority.
- Entity Data Distribution Checklist — every platform, with status, owner and recheck date.
- Verified-Tier Completion Report — before-and-after state of every profile and registry listing.
- Earned Mention Log — date, source, URL, whether linked, whether the description was accurate, and whether the source appears in the cited-domain table.
- Expert Source Kit — biographies, credentials, headshots, canonical boilerplate and topic list for each named expert, so a journalist can use them without a meeting.
- Comparison Page Register — every third-party comparison or listicle where the client should appear, with current status and accuracy.
- Paid Placement Assessment — a documented recommendation on any paid opportunity, against the two tests above.
Tools
Ahrefs or Semrush for mention monitoring and third-party comparison discovery. The Phase 01 cited-domain table as the primary targeting instrument, which is more useful for this purpose than any off-the-shelf prospecting list. Brand-monitoring alerts for unlinked mentions. Association and registry portals, most of which have no API and are worked by hand.
Honest limits. Mention-monitoring tools miss a meaningful share of unlinked mentions, particularly in PDFs, gated trade publications and closed communities. And no tool can tell you whether a particular new mention influenced a particular AI answer. The relationship between corroboration and citation is inferred from aggregate movement over months, not demonstrated per mention. We say this to clients before the work starts, not when they ask for attribution.
Success metrics
- 100% of the entity data distribution checklist complete, owned and dated.
- 100% of verified-tier listings carrying the canonical description verbatim.
- A sustained, non-spiked rate of new independent accurate mentions, agreed with the client and tracked monthly. For a mid-size firm we typically plan around three to six per month rather than a burst.
- Presence on the top-ranked domains from the cited-domain table, tracked as a simple covered-versus-total count.
- Accurate inclusion in every relevant third-party comparison page identified.
- Zero widely-cited sources asserting a materially false statement about the entity.
- Over two to three quarters: improvement in unbranded discovery share of answer, which is the prompt category corroboration most affects. Directional, and reported with the attribution caveats.
Worked example (illustrative)
Halcourt's cited-domain table showed the engines pulling most heavily, for its 84 prompts, from Health Canada's own guidance pages, two industry association sites, a standards body, one specialist regulatory news publication, and three comparison pages listing Canadian regulatory consultancies. The general business press that Halcourt's previous agency had spent two years pitching appeared in none of the sampled answers.
That single table redirected the entire budget.
Verified tier: 14 platforms, of which nine needed correction. Two association directory listings still carried the 2009 name. The Bing Places listing did not exist. The Google Business Profile listed the Calgary office at a former address.
Earned tier: the anonymized deficiency-cause dataset from Phase 03 was offered to the specialist regulatory publication, which ran a piece on it and quoted a named principal. Two of the three comparison pages had Halcourt listed with an incorrect service description; both were corrected on request, which took two emails and produced more measurable movement than the press placement. The third comparison page did not list Halcourt at all and had a documented, free inclusion process that nobody at the firm had ever completed.
One paid opportunity was declined: a "Top 20 Regulatory Consultancies in Canada 2026" listing at $4,800. The page was three months old, did not appear in the cited-domain table, and had no evidence of readership. It failed both tests.
Velocity over two quarters, illustrative: four to seven new accurate references per month, no spikes, each tied to a real event.
Common mistakes
Targeting by domain authority instead of observed citations. Budget goes to prestigious domains the engines never cite for your prompts. The fix: the cited-domain table is the target list. Prestige is a proxy; observed citation is the measurement.
Buying velocity. Forty mentions in a fortnight from similar domains with identical phrasing and no real-world trigger. The pattern is discountable at best and a risk at worst. The fix: tie every reference to a real event, vary the source types, and accept that three a month sustained beats forty once.
Treating unlinked mentions as failures. The PR report counts links, so an accurate description of the company in a heavily-cited source is scored as a miss. The fix: track mentions and links separately, and score accuracy in already-cited sources as a first-class outcome.
Ignoring the third-party comparison pages. Comparison pages dominate commercial prompts, and firms rarely check what those pages say about them. Being listed incorrectly is worse than not being listed. The fix: build the comparison page register in the first fortnight of the phase and correct the errors before pitching anything new.
Starting Phase 04 before Phase 02 is done. Mentions accumulate around an ambiguous entity and reinforce the smear. The fix: the entry criteria. Resolve the entity, then build corroboration on top of it.
Buying a listicle placement without the two tests. Money goes to a page with no readership, no editorial process and no presence in any citation sample. The fix: would the page exist unpaid, and does it appear in the cited-domain table. If both answers are no, decline.
Phase 05: Monitor, Measure & Compound
Goal
Turn the programme into a measured system that improves. Phase 05 re-runs the Phase 01 measurement on a fixed cadence, maintains the assets already earned, distinguishes real movement from engine noise, and decides what the next quarter's work should be. It is also where honesty is enforced, because this is where the temptation to over-claim lives.
Entry criteria
- Phase 01 baseline exists with a documented sampling protocol.
- Phase 03 prompt-to-URL registry exists, so movement can be traced to specific work.
- A named recipient for the monthly report who will actually act on it.
- Agreement in advance on what will be reported when numbers do not move. Deciding that after a flat month produces a different, worse report.
Inputs
- Prompt set, versioned, with any additions recorded as a new version rather than edited in place.
- Baseline and all prior sampling runs, with raw response text retained.
- Prompt-to-URL registry from Phase 03.
- Earned mention log and entity data checklist from Phase 04, with recheck dates.
- Search Console, Bing Webmaster Tools, analytics and CRM data.
- Server and CDN logs, for continued AI crawler behaviour.
What we actually do
1. Run the fixed metric set. Same definitions as Phase 01, per engine, never averaged into one headline number.
| Metric | Definition | Cadence |
|---|---|---|
| Share of answer | Runs naming the brand ÷ total runs, per engine | Monthly |
| Citation rate | Runs linking the brand ÷ total runs, per engine | Monthly |
| Prompt coverage | Prompts with at least one mention ÷ total prompts | Monthly |
| Weighted share of answer | Sales-weighted variant | Monthly |
| Competitive share | The same metrics for the competitor set | Monthly |
| Mention framing | Classified as recommended / listed / compared unfavourably / factually wrong | Monthly |
| Citation health | Earned citations still resolving ÷ total earned citations | Monthly |
| AI crawler activity | Requests per agent, status codes, pages fetched | Monthly |
| Entity consistency | Checklist rows still correct ÷ total rows | Quarterly |
| Assisted conversions | Direct and branded-search sessions following AI exposure | Quarterly |
Mention framing is the metric clients underrate. Share of answer going up while framing shifts from "recommended for X" to "a smaller alternative to Y" is not a win, and a numbers-only report will call it one.
2. Hold the sampling methodology constant. The protocol is part of the metric. Same prompts, same run count, same geography, same fresh sessions with personalization off, same day of month, same time window, same recording fields. Any change to any of these creates a discontinuity, which is recorded on the chart as a version boundary rather than quietly absorbed.
When prompts are added, and they should be as the market moves, the new prompts form v1.1 and are reported separately until they have three months of history. Blending new prompts into an existing series and reporting the combined number as a trend is the most common way this measurement gets corrupted, usually without anyone intending it.
We also record the model version wherever the engine exposes it. When a step change appears across every client in the same week, the cause is almost always a model or retrieval update, not anyone's content.
3. Audit link health seriously. Otterly.ai examined more than 20 million cited URLs across seven AI engines in a one-month snapshot and found 19.3% were dead — missing, moved or unreachable. The rate was highest on ChatGPT at 25.1% and lowest on Google AI Overviews at 12.6%. Descriptive vendor data rather than a controlled study, but the operational implication is not ambiguous.
Roughly one in five cited URLs across the ecosystem does not resolve. Some of those are yours. A citation you earned and then broke in a site migration is an asset you destroyed, and it took months to build.
So every month:
# Every URL that has ever been cited, checked for resolution
while read -r url; do
code=$(curl -s -o /dev/null -w '%{http_code}' -L --max-time 20 "$url")
hops=$(curl -s -o /dev/null -w '%{num_redirects}' -L --max-time 20 "$url")
printf '%s\t%s\t%s\n' "$code" "$hops" "$url"
done < cited-urls.txt | sort | tee link-health-$(date +%Y-%m).tsv
Anything not returning 200 is triaged: restore, redirect to the closest equivalent, or, where the content is genuinely gone, accept and record the loss. Redirect chains longer than one hop get flattened. The same check runs against every URL in the sameAs register and every third-party mention in the earned log, because those break too.
Before any migration, this list becomes a mandatory redirect map. Redirect discipline is GEO work, not an IT detail.
4. Follow the cadence table. Predictability is what makes this maintainable.
| Activity | Weekly | Monthly | Quarterly | Annually |
|---|---|---|---|---|
| Full prompt-set sampling, all engines | ✓ | |||
| Competitive benchmark on the same runs | ✓ | |||
| Citation health check | ✓ | |||
| AI crawler log review | ✓ | |||
| Mention framing review | ✓ | |||
| Earned mention log update | ✓ | |||
| Brand-mention alerts triage | ✓ | |||
| Entity consistency re-check | ✓ | |||
sameAs register verification |
✓ | |||
| Prompt set review and versioning | ✓ | |||
| Full re-baseline, Phase 01 rerun | ✓ | |||
| Content decay and rewrite audit | ✓ | |||
| Robots.txt and CDN rule re-audit | ✓ | |||
| Methodology review against vendor documentation changes | ✓ | |||
| Tool stack reassessment | ✓ |
The quarterly robots.txt re-audit exists because access regressions are common and silent. A CDN vendor change, a security review, a new WAF rule. We have seen a restored OAI-SearchBot allowance reverted within a quarter by a well-meaning security team, with nobody noticing for two months.
5. Decide what to do when numbers move. Movement needs a cause before it needs a celebration.
The sequence: is the change larger than this engine's observed month-to-month variance for this prompt set? If not, it is noise, and we say so. If yes, is it isolated to one engine or present across several? Single-engine movement points to something engine-specific, often a model or retrieval change. Cross-engine movement points to something you did. Is it concentrated in specific prompts, and do those prompts map to work shipped in the last 60 to 90 days via the registry? Did competitors move in the same direction at the same time, which usually indicates a platform change rather than a client-specific one?
Only after all four questions do we attribute. And the attribution is written as a hypothesis with its supporting evidence, not as a claim.
6. Decide what to do when numbers do not move. This is the more important case and the one most agencies handle badly.
First, distinguish "no movement" from "not enough time." Entity and structured-data work typically shows up over one to three months. Corroboration work runs on a six-to-twelve-month horizon. Parametric memory operates on training-cycle timescales that nobody outside the labs controls. Reporting failure at week six on work whose horizon is nine months is its own error.
Then work down the stack from the pillar page's layer model. Access first: are the crawlers still getting 200s, and are they actually fetching the new pages, per the logs. Then identity: did any third-party record revert. Then substance: are the new pages actually answering the sub-question, or restating the service page. Then structure: is the answer in the first 60 words. Then corroboration: has anything independent actually changed.
If all five layers are clean and three months have passed with no movement, we say so plainly and we question the prompt set. The most common cause of genuine, sustained flatness is a prompt set that measures questions the buyers do not ask. That is our error, and it gets reported as ours.
7. Report attribution honestly. There is no Search Console for generative answers. Google reports AI-feature traffic inside Search Console's existing Performance report, but there is no authoritative cross-engine view of who saw you in an AI answer and what they did next.
What we can observe: share of answer and citation rate on a fixed sample; referral traffic from AI surfaces where the engine passes a referrer, which is inconsistent; direct and branded-search volume changes; self-reported source data from lead forms, which is noisy but not worthless; and sales conversations where a buyer says an AI assistant named the firm.
What we cannot do: attribute a specific deal to a specific AI answer. We do not build models that pretend otherwise. The strongest honest statement available is usually a correlation across a quarter, presented with its confounders named. Clients respect this considerably more than a fabricated attribution model, in our experience, because they already know the fabricated ones are fabricated.
The discipline we hold to is the one Otterly.ai demonstrated in its "year in title" experiment: 11 pages, an apparent 56–61% citation increase across two waves, but one outlier page accounting for 86–93% of the gain and an untouched control page rising just as much. The authors concluded the effect could not be separated from noise. That is what a competent negative result looks like, and there is not enough of it published in this field.
flowchart TD
A["Monthly sampling run"] --> B{"Change exceeds observed variance?"}
B -->|No| C["Report as noise. No action"]
B -->|Yes| D{"One engine or several?"}
D -->|One| E["Check model or retrieval change"]
D -->|Several| F{"Maps to shipped work in registry?"}
F -->|Yes| G["Attribute as hypothesis with evidence"]
F -->|No| H{"Competitors moved too?"}
H -->|Yes| I["Platform-level change"]
H -->|No| J["Investigate access, identity, structure"]
C --> K["Quarterly re-baseline"]
E --> K
G --> K
I --> K
J --> K
K --> L["Next quarter scope"]
The decision tree we run before attributing any movement. Most months end at the first node, and saying so is the job.
Outputs & deliverables
- Monthly Visibility Report — the full metric set per engine, with the sampling protocol restated on the same page, and variance bands shown.
- Competitive Movement Summary — the client against the competitor set on identical runs.
- Citation Health Report — every cited URL checked, with a triage list.
- AI Crawler Activity Report — requests per agent, status codes, and any access regression flagged.
- Mention Framing Analysis — how the brand is characterized, with verbatim examples.
- Quarterly Re-Baseline — a full Phase 01 rerun, including prompt set review and versioning.
- Attribution Statement — what we can and cannot claim this quarter, in writing.
- Next-Quarter Scope — prioritized, traced to specific findings.
Tools
Cross-engine sampling via Ahrefs Brand Radar, Otterly.ai or Profound, depending on client coverage needs. Search Console and Bing Webmaster Tools for the search side. Log analysis for crawler behaviour. Scripted link-health checking. A spreadsheet for the metric series, because tool exports change format and a client-owned series that survives a tool change is worth more than a dashboard that does not.
Restating the limit, because it belongs in every measurement conversation: no independent third-party accuracy audit of any AI-visibility tool exists that we have been able to verify. These tools sample; they do not observe the population. Two tools will disagree on the same brand in the same week. We reconcile, report the discrepancy, and never present a vendor number as ground truth.
Success metrics
- 12 of 12 monthly sampling runs completed on protocol, with no undocumented protocol changes.
- 100% of cited URLs checked monthly; 100% of non-200s triaged within the reporting cycle.
- Zero unnoticed access regressions, measured as time between a crawler receiving a non-200 and the finding being reported.
- Prompt set reviewed and versioned quarterly, with additions reported separately until three months of history exist.
- Entity consistency checklist at 100% at each quarterly check.
- Every reported movement carrying a stated cause or an explicit statement that the cause is unknown.
- Trend direction on share of answer and citation rate against the frozen baseline. Reported as observation, never as a commitment.
Worked example (illustrative)
Halcourt's programme entered Phase 05 in month three and has been running the same 84-prompt set at five runs per engine since.
Illustrative movement, months 1 to 9:
| Engine | Baseline SoA | Month 9 SoA | Note |
|---|---|---|---|
| ChatGPT | 3.6% | 14.1% | Largest change followed the CDN 403 fix in month two |
| AI Overviews | 6.9% | 12.4% | Gradual, tracked the Phase 03 cluster publication |
| Gemini | 5.2% | 10.8% | Gradual |
| Perplexity | 9.0% | 17.6% | Step change in month three after PerplexityBot was allowed |
| Claude | 4.3% | 9.2% | Slowest and noisiest of the five |
Three findings from the reporting that mattered more than the numbers.
Month five was flat across every engine. The report said so, on the first page, and attributed it to nothing having shipped in month four because the expert review queue had stalled. No narrative was constructed.
Month six's link-health check found 11 of 63 tracked cited URLs returning 404 after a CMS upgrade silently changed the URL pattern for team pages. Consistent with the ecosystem-wide 19.3% dead-citation observation, though this instance had a specific and fixable cause. Redirects restored nine; two pages had been deleted.
Month seven showed a step change on Claude across all clients in the same week. Reported as a probable platform change, not as a result. Two months later the same prompts had drifted back down, which confirmed the reading.
Prompt set moved to v1.1 in month six, adding nine prompts about a regulatory change. Those nine were reported separately until month nine.
Common mistakes
Changing the prompt set and reporting a trend. Prompts are added or reworded, the number moves, and the movement is attributed to the work. The series is now meaningless and usually nobody notices for a year. The fix: version the set. Report additions separately until three months of history exist. Mark version boundaries on every chart.
Reporting an average across engines. The engines diverge widely, so a blended number describes no real surface and hides the one engine that collapsed. The fix: per engine, always. There is no useful single headline number.
Celebrating movement inside the variance band. Share of answer goes from 8% to 11% on five runs per prompt and the report calls it a 38% improvement. The fix: establish the variance band from the first three months and report against it. If it is inside the band, say it is inside the band.
Skipping link health because it is boring. With roughly one in five cited URLs dead across the ecosystem, an unmonitored migration quietly destroys a year of earned citations. The fix: monthly automated checking, and a mandatory redirect map before any migration.
Manufacturing attribution. A model is built that assigns revenue to AI visibility with implausible precision, and it survives exactly until a CFO examines it. The fix: state what is observable, name the confounders, and let correlation be correlation.
Treating flatness as a reporting problem. A flat month gets padded with activity metrics and secondary charts. Trust is spent to avoid an uncomfortable sentence. The fix: agree the flat-month format before month one. A flat month with a diagnosis is a better report than a good month with a story.
What this methodology will not do
This section is here because the rest of the page is a description of work we believe in, and belief is not evidence. These are real limits. They apply to us as much as to anyone else selling this service.
It will not guarantee that any AI engine names, recommends or cites you. Not for a prompt, not for a category, not ever. The ranking and selection systems are undocumented, they change without notice, and no vendor has committed to any behaviour that could be guaranteed. Any agency offering a guarantee is either misunderstanding the technology or misrepresenting it. This is not a hedge we add for legal comfort. It is the actual state of the field.
It will not produce reproducible results, because the systems are not deterministic. The same prompt, the same engine, the same minute, produces different answers. This is why our sampling protocol exists, and the protocol reduces the noise rather than removing it. A client who wants a number that does not wobble is asking for something that does not exist in this channel, and any dashboard presenting a smooth line is smoothing something.
It gives us no control over model updates. A model version change or a retrieval-stack change can move visibility in either direction, across an entire client base, in a week, for reasons no one outside the vendor can see. We can detect these events by watching for correlated movement, and we report them as what they are. We cannot predict or prevent them.
It cannot make you appear in answers that use no retrieval. When a model answers from parametric memory without searching, it draws on what it absorbed during training. That is influenced only over long horizons, by having existed prominently in the corpus that gets trained on. Every phase in this methodology contributes to that over years. None of them affect it this quarter, and a well-executed programme may not change what a model says with retrieval disabled for a long time. Level 5 on the maturity model is not purchasable.
Attribution will remain imperfect. There is no cross-engine equivalent of Search Console. Referrers are inconsistent, AI exposure often precedes a direct visit or a branded search rather than a click, and buyers do not reliably remember where they first heard a name. We report what is observable and name the confounders. We will not build an attribution model whose precision exceeds the quality of its inputs.
We cannot verify the accuracy of the tools we use. No independent third-party audit of any AI-visibility tool exists that we have been able to find. Vendors report on their own coverage with their own sampling and none publish an error rate. We use them because systematic sampling beats anecdote, and we treat their outputs as estimates.
It will not fix a weak underlying business. Generative retrieval rewards corroboration. If independent sources will not say good things about a firm, the constraint is not the marketing. Occasionally the honest finding from Phase 04 is that the corroboration does not exist because it has not been earned, and no amount of distribution work manufactures it.
It will not work without internal capacity. Expert review time, engineering release capacity, and someone able to authorize changes to third-party records. Phase 03 stalls without experts, and Phase 02 stalls without access. These are the two most common reasons a programme underdelivers, and neither is solvable by the agency alone.
Some of what we believe will turn out to be wrong. The evidence base for this field is thin. The foundational academic work predates the current model generation. Most published measurement is vendor-produced and unaudited. We have tried to be explicit about what is documented, what is inferred and what is opinion, and we will update this page when the evidence changes rather than defending a position we published in 2026.
Continue in the Knowledge Hub
- GEO pillar page — what GEO is, how the six answer surfaces differ, the maturity model, and the myths worth arguing with.
- GEO Glossary — sixty-nine defined terms including share of answer, query fan-out, entity disambiguation and citation rate.
- GEO FAQ — sixty questions, including the timeline and guarantee questions this page raises.
- GEO Resource Library — primary sources, the original GEO paper, vendor documentation and the tool comparison.
- GEO Case Studies — how these phases run against real constraints, with the trade-offs included.
Related services: AI Visibility & GEO · Search Dominance · Authority Architecture
Sector context: Legal · Healthcare · Technology · Finance
Work with Artlogic
Artlogic runs AI visibility programmes for firms expanding across Canada, the United States, Europe and the Middle East. It is Phase 01: a fixed prompt set, sampled across every surface your buyers use, benchmarked against competitors you name, with the crawler access audit attached. You will know where you stand before you decide whether to do anything about it, and you keep the prompt set either way.
{
"@context": "https://schema.org",
"@graph": [
{
"@type": "Article",
"@id": "https://artlogic.ca/geo/methodology#article",
"headline": "The Artlogic GEO Methodology",
"description": "The operational manual for Artlogic's five-phase GEO methodology: entry criteria, procedures, deliverables, tools, metrics, worked examples and the mistakes each phase prevents.",
"datePublished": "2026-08-06",
"dateModified": "2026-08-06",
"inLanguage": "en",
"isPartOf": { "@id": "https://artlogic.ca/geo#hub" },
"author": { "@id": "https://artlogic.ca/#organization" },
"publisher": { "@id": "https://artlogic.ca/#organization" },
"about": [
{ "@type": "Thing", "name": "Generative Engine Optimization" },
{ "@type": "Thing", "name": "Entity disambiguation" },
{ "@type": "Thing", "name": "Structured data" },
{ "@type": "Thing", "name": "Web crawler" }
],
"hasPart": [
{ "@type": "WebPageElement", "name": "Phase 01: AI Visibility Audit", "url": "https://artlogic.ca/geo/methodology#phase-01-ai-visibility-audit" },
{ "@type": "WebPageElement", "name": "Phase 02: Entity and Authority Foundation", "url": "https://artlogic.ca/geo/methodology#phase-02-entity-and-authority-foundation" },
{ "@type": "WebPageElement", "name": "Phase 03: Answer Engineering", "url": "https://artlogic.ca/geo/methodology#phase-03-answer-engineering" },
{ "@type": "WebPageElement", "name": "Phase 04: Citation and Consensus Building", "url": "https://artlogic.ca/geo/methodology#phase-04-citation-and-consensus-building" },
{ "@type": "WebPageElement", "name": "Phase 05: Monitor, Measure and Compound", "url": "https://artlogic.ca/geo/methodology#phase-05-monitor-measure-and-compound" }
],
"citation": [
{
"@type": "ScholarlyArticle",
"name": "GEO: Generative Engine Optimization",
"identifier": "arXiv:2311.09735",
"url": "https://arxiv.org/abs/2311.09735"
}
]
},
{
"@type": "BreadcrumbList",
"itemListElement": [
{ "@type": "ListItem", "position": 1, "name": "Home", "item": "https://artlogic.ca" },
{ "@type": "ListItem", "position": 2, "name": "GEO", "item": "https://artlogic.ca/geo" },
{ "@type": "ListItem", "position": 3, "name": "Methodology", "item": "https://artlogic.ca/geo/methodology" }
]
}
]
}
Schema note. This page describes a procedure, so
HowTolooks like the obvious type. It is not. Google removed HowTo rich results in September 2023, so the markup earns no search appearance, andHowTowould also misrepresent the content: this is a professional methodology with judgement calls and entry criteria, not a set of universally applicable steps.ArticlewithhasPartand acitationproperty does more useful work. It makes the page's structure and its sourcing machine-readable, which is the point. There is deliberately noFAQPageblock either, for the reason set out in Phase 03: Google's documentation states FAQ rich results stopped appearing on 7 May 2026, with support withdrawn through June and August 2026.
Last reviewed 6 August 2026. Halcourt Regulatory Partners is an illustrative example invented for this page and is not a client; every figure attached to it is labelled illustrative. Sourced figures carry their source on the line where they appear, and vendor-produced data is identified as such.