Published
Author

Callum
Categories
AEO & GEO

Ready to take it a step further?
Let's start talking about your project or idea and find out how Creato can help your business grow.
Technical AEO and GEO work starts with a practical question: can a search system retrieve accurate, usable evidence from your website when someone asks a relevant question?
A page can look excellent in a browser and still return an empty application shell to a crawler. It can be indexed but restricted from supplying snippets. It can contain the right information while separating a price from the conditions that make that price meaningful. Each problem needs a different fix.
This guide follows the path from discovery to citation. It is written for developers, technical marketers and business owners commissioning AEO and GEO services. The recommendations distinguish documented platform behaviour from implementation choices worth testing. The examples are illustrative; they do not describe Creato client results.
For the broader website brief, start with our AEO website design guide. Here, we will examine the machinery underneath it.
1. Understand the stages between a web page and an AI answer
Retrieval-augmented generation, or RAG, combines a language model with retrieved information. The original RAG research paper describes a system that retrieves passages from an external index and conditions generation on them. This is a useful architectural model, although it does not reveal the complete implementation of any current commercial search product.
For diagnosis, separate six stages:
Discovery: a system learns that a URL exists.
Fetching: it requests the resource and receives a response.
Processing: it extracts usable content, potentially after rendering.
Indexing: a search system stores representations it can retrieve.
Retrieval: a query selects candidate documents or passages.
Answering and attribution: a system uses selected information and may display supporting links.
These are diagnostic categories, not a claim that every assistant implements the same pipeline. Products can combine search providers, cached content and live fetches. An assistant can also answer without searching.
Google documents that AI Overviews and AI Mode may issue related searches across subtopics through query fan-out. A supporting page must be indexed and eligible for a Search snippet. There is no additional technical eligibility requirement. See Google's AI feature guidance.
Consider a question about choosing a Sydney architect for a heritage renovation. Useful supporting material might cover heritage experience, location, project constraints and the appointment process. One service page cannot necessarily answer every subquestion. Map the evidence your site actually provides before deciding what to publish.
2. Separate search crawling, model training and user-requested visits
“Allow AI” is an imprecise technical requirement. Record the intended product and purpose before changing access controls.
OpenAI's three relevant user agents
OpenAI's crawler documentation distinguishes:
OAI-SearchBot: supports discovery in ChatGPT search. Blocking it excludes a site from search answers, although navigational links can still appear.
GPTBot: crawls material that may be used to train OpenAI's foundation models.
ChatGPT-User: handles certain user-initiated visits. It is not the automatic search crawler, does not determine search inclusion, and robots.txt rules may not apply to these user-initiated actions.
Search and training permissions are independent. A publisher choosing search access while declining training crawling could use these groups:
This is a policy example, not a replacement robots.txt file. Preserve restrictions needed for other paths and agents. Review the complete configuration before deploying it.
Google's controls have different scopes
Googlebot is relevant to Google Search. Google-Extended controls certain Gemini training and grounding uses; Google states that it does not affect inclusion or ranking in Google Search. It has no separate HTTP request user-agent string.
Keep a small access register: product, agent or token, allowed paths, hosting restrictions, responsible owner and reason for the decision. That prevents a later security update from silently reversing the agreed publishing policy.
Robots rules do not authenticate visitors or protect confidential files. Put private content behind appropriate access controls. Never publish secrets and rely on a crawler preference to keep them private.
3. Diagnose indexing and preview controls independently
A successful response is only one check. Google's technical requirements include crawler access, an HTTP 200 response and indexable content. Meeting those conditions establishes eligibility, not guaranteed indexing.
Robots.txt and noindex do different jobs
A robots.txt disallow rule restricts crawling. A noindex directive asks a supporting search engine to exclude the resource from its index. Google must fetch a page to see its noindex meta tag or X-Robots-Tag header. Blocking that fetch can leave the instruction undiscovered; a URL can still appear based on other information. Google does not support putting noindex in robots.txt. See its noindex documentation.
For an intended public landing page, inspect both the HTML and response headers. A CMS setting may remove the HTML directive while a hosting rule continues adding a noindex header. A visual review will not reveal that conflict.
Snippet restrictions can affect answer use
Google documents that nosnippet prevents content being used as direct input for AI Overviews and AI Mode. Max-snippet limits the amount that may be used, with documented exceptions for separately granted permissions. Data-nosnippet can exclude selected text from snippets. Audit these deliberately, following the robots meta and preview specifications; do not remove a publisher's intended restrictions simply to chase visibility.
Make the preferred URL unambiguous
For duplicate or substantially similar URLs, Google treats redirects and rel="canonical" as strong canonicalisation signals; sitemap inclusion is weaker. They express a preference rather than absolute control. Its canonical URL guidance explains the interaction.
Check that the canonical, sitemap and internal links agree. During a redesign, a copied template can accidentally point every article at the same canonical URL. Keep distinct service pages self-consistent; do not canonicalise genuinely different material merely because it targets related questions.
4. Inspect the response, rendered page and crawler evidence
The browser view is one observation. The initial HTTP response and a search engine's processed version provide different evidence.
Google can render JavaScript, but crawling, rendering and indexing are separate stages. Blocked resources can prevent rendering, and Google recommends server-side or pre-rendering because not all bots execute JavaScript. These behaviours are documented in its JavaScript SEO guidance.
A compact fetch check
On a computer with curl, replace the example URL with a page you manage:
This makes a GET request, follows redirects, saves the response body and records the response headers. Review every redirect and the final status. Search the saved HTML for a distinctive sentence from the main content, the canonical URL and robots directives.
Finding a sentence somewhere in the file is not sufficient: it might exist only inside a script or data payload. Confirm that essential content is ordinary document text in the rendered DOM, and compare its meaning with the initial response. A 200 response containing only an access challenge is also a failed content fetch.
Use three separate evidence sources
Direct inspection: response status, headers, redirects and meaningful HTML.
Search tooling: Google's URL Inspection tool distinguishes indexed information from a live test. A successful live test does not prove indexing or inspect every indexing condition.
Server and CDN logs: requested URL, timestamp, status, bytes served and any security challenge. Look for repeated failures on important pages.
A user-agent string can be copied. For Google requests, use its published verification methods rather than trusting a bot name in a log. For other providers, follow their current identity and IP guidance.
Prefer a narrowly scoped correction to a confirmed blocking rule over disabling the firewall. Retest the affected URL and normal security behaviour. A successful Google inspection still says nothing conclusive about another provider's access.
5. Understand retrieval beyond keyword matching
A retrieval system may combine several methods. Lexical retrieval matches terms; vector retrieval compares numerical representations of meaning; a reranker can reconsider the candidates against the query.
Azure AI Search's hybrid search documentation gives a concrete implementation: it combines full-text and vector retrieval. Exact names, specialised terms and identifiers can benefit from lexical matching, while vector search can recover conceptually related language.
This explains a useful editorial balance. Name the actual service, place, organisation and technical terms accurately. Explain them in natural language that resolves the customer's problem. Repeating an exact phrase throughout every heading adds little useful information.
A worked ranking example
Azure AI Search documents reciprocal rank fusion, which combines ranked lists using contributions of 1 / (k + rank). With an illustrative constant of 60:
Document A ranks first in one list and twentieth in another: 1/61 + 1/80, approximately 0.02889.
Document B ranks fourth in both lists: 1/64 + 1/64, or 0.03125.
Document B wins this simplified calculation. The example demonstrates why retrieval can depend on multiple signals rather than one apparent keyword position. It is not a formula for ranking a website in ChatGPT or Google. Azure's published implementation does not establish those products' undisclosed scoring methods.
6. Write passages that preserve their own meaning
Document processing can divide a long page into smaller chunks. Microsoft's chunking guidance describes fixed-size, structure-based and semantic approaches. Chunking can address model input limits and improve how mixed-topic material is represented.
A website publisher usually cannot choose how an external search system splits a page. The practical response is to make important sections understandable when read independently. This is an information-design recommendation, not a documented universal ranking factor.
Keep the claim and its conditions together
Consider this hypothetical service description:
Ambiguous: “It starts at $4,000. Everything is included. Contact us to discuss the details.”
More useful: “Example Studio's five-page website package starts at AUD $4,000 excluding GST. It includes design and development using client-supplied copy. Copywriting, paid integrations and ongoing hosting are quoted separately.”
The second passage identifies the provider, deliverable, currency, tax treatment, inclusions and exclusions together. It does not force a reader to recover those qualifications from a distant footer. The amount is fictional and should never be reused as actual pricing.
Apply the same test to project evidence. Place the result beside the measurement period, sample and limitation. “Enquiries increased” is incomplete without a baseline and a definition of an enquiry.
Use structure that survives extraction
Write descriptive H2 and H3 headings that identify the question or subject.
Introduce the organisation or service by name before relying on pronouns.
Keep units, dates, geography and important exceptions close to the relevant claim.
Use real lists for steps and genuine tables only when relationships need them.
Provide explanatory text for information otherwise available only in graphics or video.
A useful editorial check is to copy one section into a blank document. Can a colleague explain what it means without the rest of the page? Repair missing context rather than forcing every answer into a supposed ideal word count. No source cited here establishes a universal passage length for AI search.
7. Model entities and structured data without inventing authority
Entity clarity means making it possible to distinguish your organisation, people, locations and services from similarly named things. It starts with visible facts: a consistent business name, an accurate service description, genuine locations and identifiable authors.
Google's Organization documentation describes properties including name, URL, logo and sameAs links. Use sameAs for genuine profiles of that organisation, not unrelated articles or prestigious sites you want to be associated with.
For an article, Article or BlogPosting markup can describe the headline, author, image and publication or modification dates. Connect the author to an appropriate profile where one exists. Dates should reflect actual publication and meaningful changes.
Developers can use a stable identifier such as a canonical domain followed by #organization to reference the same entity across a site's JSON-LD graph. Treat that as a maintainable modelling convention. It does not force an external knowledge graph to merge an entity or confer authority.
Check three things before release:
The JSON parses and the types and properties are appropriate.
The factual values agree with what visitors see.
There are no conflicting duplicate graphs from the CMS, plugins or custom code.
Google's structured data policies require relevant, representative markup. Technical validity alone does not establish eligibility for a rich result, and a rich result is not an AI citation guarantee.
Google's AI feature guidance also says these experiences require no special schema or AI text file. An experimental llms.txt file should therefore have a defined use case and evaluation plan, rather than replacing the access and content work described above.
8. Make evidence auditable at the claim level
A citation is useful only if its source supports the nearby claim. A page can be cited while the generated answer still omits an important condition or attributes a statement incorrectly.
The ALCE research benchmark separates answer correctness from citation quality. That distinction is valuable for website audits: check the accuracy of the statement and whether the linked source supports it. Its historical benchmark results should not be treated as present-day failure rates for commercial assistants.
Maintain an evidence record for commercially important claims:
The precise claim and the page that publishes it.
The primary source, permission or underlying records.
The measurement method and dates, if the claim is quantitative.
The limitations a reader needs to interpret it correctly.
The person responsible for review when circumstances change.
When publishing original research, explain how the sample was chosen, how observations were classified and what the study cannot show. Publish enough methodology for a reader to assess the result without exposing private client data.
The GEO research paper evaluates content interventions within a defined experimental framework. Such studies can motivate hypotheses. They do not establish that adding statistics or citations mechanically increases visibility across every current product. Add evidence because it supports a useful claim, then measure what happens.
9. Keep discovery and freshness signals consistent
Useful content needs a discoverable URL. Google recommends crawlable links using anchor elements with href attributes, with descriptive link text that provides context. Follow its link guidance when connecting guides, services and relevant project examples.
Give each page a job. A technical guide can answer implementation questions and link to the service that provides help. An audit guide can explain measurement. Avoid several nearly identical pages competing to answer the same question.
A sitemap helps communicate preferred URLs. Its lastmod values should reflect significant page changes, not a date refreshed automatically on every request.
For participating search engines, IndexNow can notify them when content is added, changed or removed. Submission does not guarantee indexing, immediate processing or AI citations. Preserve evidence of successful publication and discovery separately from evidence of later inclusion.
10. Measure each stage with the right denominator
An access check, a citation and a qualified enquiry answer different questions. Keep them separate in reporting.
Google includes AI Overviews and AI Mode activity within Search Console's overall Web performance data. That aggregate cannot be labelled as a standalone AI traffic series.
Bing's AI Performance reporting covers supported Microsoft and partner experiences. Its June 2026 preview expansion adds intents, topics, citation share and comparison views. Citation share is an observational proportion of citations for a grounding query, not traffic share or a content quality score. Check the coverage and definitions before combining it with other measurements.
Build a reproducible observation record
Our GEO audit checklist covers the baseline process. For technical experiments, add the release identifier and evidence of recrawling where available. Record:
Exact prompt, platform, visible model or mode, date and location context.
Whether search ran and whether an AI answer appeared.
Business mention, recommendation and own-domain citation as separate fields.
Cited URL, supported claim and factual errors.
Valid answers, unavailable responses and tool failures separately.
Suppose an illustrative run produces 40 valid answers and eight cite your domain. That is a 20% citation rate within that sample. It is not a share of all customer searches. Do not quietly remove difficult prompts or failed runs from the record.
Separate an observed improvement from its cause
Define the hypothesis before editing: “Making exclusions explicit may reduce incorrect package descriptions.” Change the relevant passage, preserve a dated copy and repeat the same tests after the updated content is observable.
Where practical, compare similar unchanged pages and retain the same question set. Even then, platform updates, competing content and personalisation can affect outcomes. Report the observation and remaining uncertainty. “The updated page was cited in this answer” is supportable; “this heading caused the model to recommend us” usually requires much stronger evidence.
11. Turn the audit into release criteria
Write implementation tickets with a specific failing condition and a way to verify the fix. “Improve AI readiness” leaves both the developer and the business owner guessing.
For a priority public service page, a practical acceptance checklist is:
Access: the intended URL resolves to the expected content without authentication or an unintended security challenge.
Directives: crawler rules, headers and HTML reflect the agreed publishing policy.
Rendering: essential facts appear as readable text and remain coherent on mobile.
Identity: redirects, canonical, internal links and sitemap use the intended URL consistently.
Content: scope, geography, exclusions and evidence are explicit enough to interpret independently.
Markup: structured data parses, matches the page and contains no unsupported claims.
Observation: inspection results and baseline measurements are saved with the release date.
Prioritise confirmed failures before speculative experiments. An accidental noindex on a core page has a clear mechanism and a verifiable repair. One missing mention in an AI answer has many possible explanations.
The technical goal is a website that exposes accurate, well-supported information reliably, with measurements that tell you which stage needs attention. Selection by an external answer system remains outside a publisher's control.
If you want help applying this process to your business, explore Creato's AEO and GEO services in Sydney and Australia. Start with the pages and customer questions that matter commercially, then turn the findings into specific technical and content improvements.

