Start with the product and the task
“AI search” covers several experiences. A Google result with a generated answer, a ChatGPT conversation using web search and a Perplexity answer are not one shared ranking system. Record the exact product and mode you are checking. Avoid assuming that a position in one search engine determines every answer elsewhere.
Begin with a relevant page and a real question from your audience. Keep the intended next step visible: understanding a service, comparing a scope or making an enquiry.
Check the page before changing crawler rules
- Open the canonical URL and verify the response status.
- Check robots meta directives and HTTP indexing headers.
- Review the exact path in robots.txt, including named groups and the longest matching rule.
- Use Search Console or the relevant webmaster tool to inspect actual indexing where available.
- Check hosting logs and security rules if a legitimate crawler appears to be challenged.
Keep intentional private areas restricted. Robots.txt is a crawl preference, not an access-control system for confidential content.
ChatGPT and OpenAI crawlers
OpenAI documents OAI-SearchBot for search, GPTBot for potential training use and ChatGPT-User for user-initiated requests. Configure these separately. Review OpenAI’s crawler documentation for current details before editing a rule.
A user-agent string alone does not prove a request came from OpenAI. Use the provider’s verification information when investigating logs. Check whether the content is publicly available and useful before attributing a missing citation to a robots rule.
Google Search and AI Overviews
Review Googlebot access and the page’s status in Search Console. Google’s AI features documentation distinguishes Search access from Google-Extended controls. A training preference should not be mistaken for the control that determines ordinary Google Search crawling.
Check that the page answers the relevant question and supports its claims. There is no technical switch that instructs an AI Overview to cite a particular business.
Perplexity
Perplexity documents PerplexityBot for search and Perplexity-User for user requests. Its crawler documentation explains their roles and verification information. Review the relevant rule and the actual page response.
Record the sources in an observed answer rather than assuming a particular type of page will always be preferred. A source link proves what was shown in that sample, not why it was selected.
Gemini
Identify whether you are observing Gemini as a standalone product or a generated feature within Google Search. Record the product and mode in your review. Read the applicable Google documentation before changing access controls; one observation should not be generalised to every Google AI experience.
Microsoft Copilot and Bing
Use Bing Webmaster Tools to review Bing discovery and indexing, and inspect the current Bing webmaster guidelines. Treat a Bing result and a Copilot answer as distinct observations. Indexing is useful to verify, but it does not establish that a particular answer will cite you.
Make the destination useful
Describe the service clearly, explain the scope and link to relevant evidence. For a practice, show accurate clinician and location information with responsible clinical review. For a charity, explain the cause, governance and donation journey. These are practical editorial priorities, not a published list of model ranking factors.
Check structured data against the page
Use appropriate organisation, service, article or breadcrumb data where it helps describe the content. Validate its syntax and compare every material claim with what a visitor can see. Markup does not replace the underlying information or prove eligibility for a search feature.
Keep a dated record
Record the query, date, product, relevant account or location context, answer and cited URLs. Separate a brand mention from a citation and both from a website visit. Review search reports and received enquiries alongside the observations.
Our measurement guide explains the distinctions. Our robots checker checks declared rules only; it cannot confirm indexing or firewall access.