AI SEARCH CRAWLER
REFERENCE.

Which crawlers and fetchers each AI search provider documents, what the provider says each one is for, and how to check that a request claiming to be one of them is genuine. Every row links the provider’s own document, carries one of five classes, and shows the date that document was last read.

This page describes controls. It does not recommend allowing or blocking anything — what to do with these controls depends on your content and your business — and no provider on this page guarantees that an allowed crawler leads to a citation.

HOW TO READ THIS PAGE

Five classes · a verified date on every row · a re-verification date

Every row on this page carries three things: a link to the provider’s own document, one of five classes, and the date that document was last read. The classes are the same five we use on every claim we make — defined once, on the methodology page.

DOCUMENTED

Directly supported by the platform provider’s own documentation.

OBSERVED

Seen in our testing, but not established as a platform rule.

INFERRED

A reasonable reading of the evidence. Not a mechanism.

UNVERIFIED

Insufficient evidence. Says so, and stays unused.

EXPERIMENTAL

A deliberate test of something unverified, registered before it runs.

On this page a row is DOCUMENTED only when the provider states it in its own documentation, linked and dated. A crawler that appears in server logs but in no provider document is OBSERVED, never DOCUMENTED. Where we draw a conclusion from documented facts, that conclusion is marked INFERRED, separately from the facts. Nothing is promoted a class because it is convenient.

Verified
the date we last read the row’s source. Every row carries its own; today they all read .
Next re-verification
— every source re-read, every moved or changed source marked and re-dated.
Quotes
always in the provider’s English, marked as quotations, in both language versions. Translating a provider’s wording would turn DOCUMENTED into INFERRED.

How we label a claim — on the methodology page

VERIFYING THAT A REQUEST IS GENUINE

Published endpoints · reverse DNS · the layer in front of the origin

A user-agent string is a claim, not an identity. Anyone can send one. What a provider publishes to settle the question is an IP list, a reverse-DNS pattern or a tool — and every endpoint below answered on the date in its row.

ProviderMethodPublished endpointsCheckedClassVerifiedSource
OpenAI Published IP ranges, one list per agent. HTTP 200 on DOCUMENTED OpenAI · crawler documentation
Anthropic If a crawler has a source IP address on this list, it indicates that the crawler is coming from Anthropic. One published IP list for all Anthropic agents. HTTP 200 on DOCUMENTED Anthropic · does Anthropic crawl data from the web
Perplexity Published IP ranges, one list per agent. HTTP 200 on DOCUMENTED Perplexity · crawlers
Google The common crawlers generally crawl from the IP ranges published in the common-crawlers.json object, and the reverse DNS mask of their hostname matches crawl-***-***-***-***.googlebot.com or geo-crawl-***-***-***-***.geo.googlebot.com. Published IP ranges per crawler group, plus a reverse-DNS mask for the common crawlers. HTTP 200 on DOCUMENTED Google · common crawlers
Apple Traffic coming from Applebot is generally identified by using reverse DNS in the *.applebot.apple.com domain. Another way is to match the IP address with a CIDR prefix contained in the following JSON file Reverse DNS in *.applebot.apple.com, or a published IP list. HTTP 200 on DOCUMENTED Apple · about Applebot
Mistral Published IP lists for the index and user agents. HTTP 200 on DOCUMENTED Mistral · robots
Amazon Published IP addresses, one page per agent. HTTP 200 on DOCUMENTED Amazon · Amazonbot
Microsoft The Verify Bingbot tool inside Bing Webmaster Tools. a tool behind a login — nothing to fetch DOCUMENTED Bing · which crawlers does Bing use transcribed from the rendered page — the table is JavaScript-rendered

Two things the documentation adds

Alternate methods like blocking IP address(es) from which Anthropic Bots operates may not work correctly or persistently guarantee an opt-out, as doing so impedes our ability to read your robots.txt file. Anthropic · DOCUMENTED · 2026-09-05

The IP list is published for recognising Anthropic’s crawlers. Anthropic states that blocking those addresses is not a reliable way to opt out — robots.txt is the documented control.

If you’re using a Web Application Firewall (WAF) to protect your site, you may need to explicitly whitelist Perplexity’s bots to ensure they can access your content. Perplexity · DOCUMENTED · 2026-09-05

The check may have to happen in front of the origin. A robots.txt that permits a crawler does not help if a firewall or bot-protection layer answers that crawler with a challenge — a failure that is invisible in robots.txt and invisible in analytics. Perplexity’s documentation says so in as many words and lists firewall settings for it.

THE THREE JOBS A CRAWLER DOES

Search · user retrieval · training

The tables on this page are sorted by what an agent does, because that is what the controls attach to. The same provider often runs one of each. This three-way split is our framing — INFERRED from the documented purposes, not a taxonomy any provider publishes — and the rows below say, per provider, where the documentation actually separates the three and where it does not.

JobWhat it doesWhat blocking it costs
Search Builds an index so the provider can surface and link your pages in its search product. Your pages stop being eligible for that product’s answers.
User retrieval Fetches one page because a person asked a question that needs it, right now. The assistant cannot open your page for a user who asked about it.
Training Collects content that may be used to train foundation models. Your content is excluded from future training sets — with the providers that document the separation.

SEARCH BOTS

The agent to allow if you want to appear in that provider’s answers

ProviderUser agentProvider’s own wordsClassVerifiedSource
OpenAI OAI-SearchBot OAI-SearchBot is for search. OAI-SearchBot is used to surface websites in search results in ChatGPT’s search features. Sites that are opted out of OAI-SearchBot will not be shown in ChatGPT search answers, though can still appear as navigational links. DOCUMENTED OpenAI · crawler documentation
Google Googlebot Crawling preferences addressed to the Googlebot user agent affect Google Search (including Discover and all Google Search features), as well as other products such as Google Images, Google Video, Google News, and Discover. DOCUMENTED Google · common crawlers
Microsoft bingbot our standard crawler … handles most of our crawling needs each day transcribed from the rendered page — the table is JavaScript-rendered DOCUMENTED Bing · which crawlers does Bing use transcribed from the rendered page — the table is JavaScript-rendered
Perplexity PerplexityBot PerplexityBot is designed to surface and link websites in search results on Perplexity. It is not used to crawl content for AI foundation models. DOCUMENTED Perplexity · crawlers
Anthropic Claude-SearchBot Claude-SearchBot navigates the web to improve search result quality for users. It analyzes online content specifically to enhance the relevance and accuracy of search responses. Disabling Claude-SearchBot on your site prevents our system from indexing your content for search optimization, which may reduce your site’s visibility and accuracy in user search results. DOCUMENTED Anthropic · does Anthropic crawl data from the web

OpenAI also documents a fourth agent, OAI-AdsBot, which belongs to none of the three classes: OAI-AdsBot only visits pages submitted as ads, and the data collected by OAI-AdsBot is not used to train generative AI foundation models. — DOCUMENTED · 2026-09-05 · OpenAI · crawler documentation.

Two entries that are absences

Google documents no separate crawler for AI Overviews or AI Mode.

To be eligible to be shown in generative AI features on Google Search, a page must be indexed and eligible to be shown in Google Search with a snippet, fulfilling the Search technical requirements. Just because a page meets all requirements, best practices, and complies with the policies, doesn’t mean that Google will crawl, index, or serve its content. Indexing and serving aren’t guaranteed. Google · DOCUMENTED · 2026-09-05

Google’s list of common crawlers contains no agent specific to its generative AI features, and its guide ties eligibility for those features to Search eligibility. Our reading: there is nothing extra to allow for Google’s AI answers — Googlebot is the control, and the documented extra condition sits in Search Console, not in robots.txt (see the page-level controls below). INFERRED · Google · AI features and your website

Microsoft documents no separate Copilot crawler.

Bing’s documented crawlers are bingbot, AdIdxBot, BingPreview, MicrosoftPreview and BingVideoPreview — read from the rendered page, because the table is JavaScript-rendered. None is specific to Copilot. Our reading: the controls Microsoft documents for AI answers are page-level meta values, not a user agent (see the page-level controls below). INFERRED · Bing · which crawlers does Bing use · transcribed from the rendered page

USER-RETRIEVAL AGENTS

One page, because a person asked — and robots.txt may not apply

These agents fetch a page because someone asked a question that needs it. Of the providers on this page that document one, five state a user-initiated exception to robots.txt — OpenAI, Perplexity, Google, Meta and Amazon, the last two in the five-provider table further down. Two document such an agent without stating an exception: Anthropic, which states blanket compliance, and Mistral. Whether that difference shows in practice is not something a table of documentation can answer — that would be OBSERVED, and it would need server logs. It is recorded as an asymmetry in the documentation, nothing more.

ProviderUser agentProvider’s own wordsrobots.txtClassVerifiedSource
OpenAI ChatGPT-User When users ask ChatGPT or a CustomGPT a question, it may visit a web page with a ChatGPT-User agent. […] ChatGPT-User is not used for crawling the web in an automatic fashion. Because these actions are initiated by a user, robots.txt rules may not apply. DOCUMENTED OpenAI · crawler documentation
Perplexity Perplexity-User Perplexity-User supports user actions within Perplexity. When users ask Perplexity a question, it might visit a web page to help provide an accurate answer and include a link to the page in its response. […] It is not used for web crawling or to collect content for training AI foundation models. Since a user requested the fetch, this fetcher generally ignores robots.txt rules. DOCUMENTED Perplexity · crawlers
Anthropic Claude-User Claude-User supports Claude AI users. When individuals ask questions to Claude, it may access websites using a Claude-User agent. Claude-User allows site owners to control which sites can be accessed through these user-initiated requests. Anthropic’s Bots respect “do not crawl” signals by honoring industry standard directives in robots.txt. No user-initiated exception is stated. DOCUMENTED Anthropic · does Anthropic crawl data from the web
Google Google-Agent · Google-GeminiNotebook · Google-Read-Aloud · Google-Pinpoint · … Google-Agent is used by agents hosted on Google infrastructure to navigate the web and perform actions upon user request. The Gemini Notebook fetcher requests individual URLs that Gemini Notebook users have provided as sources for their projects. Because the fetch was requested by a user, these fetchers generally ignore robots.txt rules. DOCUMENTED Google · user-triggered fetchers
  • Not a permission. “May not apply” describes the provider’s behaviour. It does not make the fetch authorised, and it does not affect any other legal or contractual control you have.
  • Not a search route. OpenAI: ChatGPT-User is not used to determine whether content may appear in Search. Please use OAI-SearchBot in robots.txt for managing Search opt outs and automatic crawl. Allowing ChatGPT-User does not get a page into ChatGPT search.

TRAINING BOTS

A data-licensing decision — separate from search where the provider documents it as separate

Blocking these is a decision about your content. With the three providers below, the documentation separates that decision from search visibility — twice in the provider’s own words, once as our inference — and the column says which is which.

ProviderUser agentProvider’s own wordsBlocking it costs search visibility?ClassVerifiedSource
OpenAI GPTBot GPTBot is used to make our generative AI foundation models more useful and safe. It is used to crawl content that may be used in training our generative AI foundation models. a webmaster can allow OAI-SearchBot in order to appear in search results while disallowing GPTBot to indicate that crawled content should not be used for training OpenAI’s generative AI foundation models. DOCUMENTED DOCUMENTED OpenAI · crawler documentation
Anthropic ClaudeBot ClaudeBot helps enhance the utility and safety of our generative AI models by collecting web content that could potentially contribute to their training. When a site restricts ClaudeBot access, it signals that the site’s future materials should be excluded from our AI model training datasets. A separate agent from Claude-SearchBot, each with its own documented effect. That blocking one leaves the other untouched is not stated in those words. INFERRED DOCUMENTED Anthropic · does Anthropic crawl data from the web
Google Google-Extended robots.txt token — no user-agent string of its own Google-Extended is a standalone product token that web publishers can use to manage whether content Google crawls from their sites may be used for training future generations of Gemini models that power Gemini Apps and Vertex AI API for Gemini and for grounding […] in Gemini Apps and Grounding with Google Search on Vertex AI. Google-Extended doesn’t have a separate HTTP request user agent string. Crawling is done with existing Google user agent strings; the robots.txt user-agent token is used in a control capacity. Google-Extended does not impact a site’s inclusion in Google Search nor is it used as a ranking signal in Google Search. DOCUMENTED DOCUMENTED Google · common crawlers

The most misread control on this page

Google-Extended does not govern AI Overviews or AI Mode. Those are Search, and Search is governed by Googlebot. Google-Extended governs Gemini model training and grounding in Gemini Apps and Vertex AI, and Google states that it does not affect a site’s inclusion in Search. That a site which disallows Google-Extended can still appear in Google’s AI answers follows from those two documented statements — INFERRED, not quoted.

Separate controls do not always mean separate crawls. OpenAI: If your site has allowed both bots, we may use the results from just one crawl for both use cases to avoid duplicative crawling. — DOCUMENTED · 2026-09-05.

CONTROLS THAT ARE NOT USER AGENTS

Page level: a meta value, a Search Console setting

Not every control is a line in robots.txt. These work at page level, without any user agent, and four providers document one. Microsoft’s two tags separate three states: in the AI answer with content (no tag), in the answer as URL, title and snippet only (NOCACHE), out of the answer entirely (NOARCHIVE) — and Microsoft states that both tags keep the page in ordinary search results.

ControlProviderProvider’s own wordsClassVerifiedSource
NOCACHE (robots meta) Microsoft Content with the NOCACHE tag may be included in Bing Chat answers. We will only display URL/Snippet/Title in the answer; Going forward, for content in our Bing Index that is labeled NOCACHE, only URLs, Titles and Snippets may be used in training Microsoft’s generative AI foundation models. Blog post of September 2023; the product is called “Bing Chat” there. DOCUMENTED Bing Webmaster Blog · NOCACHE and NOARCHIVE
NOARCHIVE (robots meta) Microsoft Content tagged NOARCHIVE will not be included in Bing Chat answers, not be linked to in the answers. Going forward, for content in our Bing Index that is labeled NOARCHIVE, we will not use the content for training Microsoft’s generative AI foundation models. We can assure publishers that content with the NOCACHE tag or NOARCHIVE tag will still appear in our search results.Blog post of September 2023; the product is called “Bing Chat” there. DOCUMENTED Bing Webmaster Blog · NOCACHE and NOARCHIVE
Search Console · “Search generative AI features” Google In addition to the technical requirements for Search, a site must be included in Search generative AI features in Search Console to be eligible for display in generative AI features on Google Search. DOCUMENTED Google · AI features and your website
nosnippet (robots meta) Apple Apple will not use data tagged nosnippet as additional context and up-to-date content when AI models are used to generate output for display in Apple products and services. Even if you disallow Applebot-Extended and tag website content with the nosnippet meta tag, your website instructions may still allow Applebot to crawl your webpages. DOCUMENTED Apple · about Applebot
noarchive (robots meta) Amazon When these user agents access web pages they respect the link-level rel=nofollow directive, and page level robots meta tags of noarchive (do not use the page for model training), noindex (do not index the page) and none (do not index the page). DOCUMENTED Amazon · Amazonbot

FIVE MORE PROVIDERS

Mistral · Meta · Apple · Amazon · Common Crawl

Five providers that document their agents and are not in the tables above. Mistral documents a clean triple. Meta documents a search-class agent, a user fetcher and a training crawler — read from the rendered page, because the raw fetch was refused, and marked as such in every row. Apple’s search crawler also feeds its model training unless a second token is disallowed. Amazon documents three agents with independent settings. Common Crawl is neither a search engine nor an assistant.

ProviderUser agentJobProvider’s own wordsrobots.txtClassVerifiedSource
Mistral MistralAI-Index Search MistralAI-Index is for automated crawling of the web for indexing purposes only. It indexes content for Mistral search, which helps answer user questions in Vibe. Content crawled by MistralAI-Index is not used for generative AI training of any kind. DOCUMENTED Mistral · robots
Mistral MistralAI-User User retrieval MistralAI-User is for user actions in Vibe. When users ask Vibe a question, it may visit a web page to help answer and include a link to the source in its response. […] It is not used for crawling the web in any automatic fashion, nor to crawl content for generative AI training. DOCUMENTED Mistral · robots
Mistral MistralAI-Training Training MistralAI-Training crawls web content to help build datasets for training Mistral generative AI models. […] This crawler is not used for search indexing or to answer live user queries in Vibe. DOCUMENTED Mistral · robots
Meta Meta-WebIndexer Search navigates the web to improve Meta AI search result quality cite and link to your content in Meta AI’s responsestranscribed from the rendered page — the raw fetch was refused DOCUMENTED Meta · web crawlers transcribed from the rendered page — the raw fetch was refused
Meta Meta-ExternalFetcher User retrieval fetches individual links at a user’s request transcribed from the rendered page — the raw fetch was refused might bypass robots.txt transcribed from the rendered page — the raw fetch was refused DOCUMENTED Meta · web crawlers transcribed from the rendered page — the raw fetch was refused
Meta Meta-ExternalAgent Training crawls the web for use cases such as training foundation AI models or improving products transcribed from the rendered page — the raw fetch was refused DOCUMENTED Meta · web crawlers transcribed from the rendered page — the raw fetch was refused
Apple Applebot Search The data crawled by Applebot is used to power various features, such as the search technology integrated into many user experiences in Apple’s ecosystem including Spotlight, Siri, and Safari. The data crawled by Applebot may also be used to help train Apple foundation models powering generative AI features across Apple products […]. Web publishers can opt-out from having their content used to train generative foundation models by disallowing Applebot-Extended in the robots.txt file. DOCUMENTED Apple · about Applebot
Apple Applebot-Extended robots.txt token — no user-agent string of its own Training Applebot-Extended does not crawl webpages. Webpages that disallow Applebot-Extended can still be included in search results. Applebot-Extended is only used to determine how to use the data crawled by the Applebot user agent. DOCUMENTED Apple · about Applebot
Amazon Amzn-SearchBot Search Amzn-SearchBot is used to improve search experiences in Amazon products and services. By permitting Amzn-SearchBot access to your website, your content is eligible to appear in search experiences such as Alexa. If robots.txt files don’t mention Amzn-SearchBot but allow other search bots, Amzn-SearchBot will crawl in accordance with the robots.txt directives given to other search bots. Amzn-SearchBot does not crawl content for generative AI model training. DOCUMENTED Amazon · Amazonbot
Amazon Amzn-User User retrieval Amzn-User supports user actions, such as responding to Alexa queries that require up-to-date information. […] Amzn-User does not crawl content for generative AI model training. Because actions taken by Amzn-User can be initiated by a user, it may not follow all robots.txt directives. DOCUMENTED Amazon · Amazonbot
Amazon Amazonbot Training Amazonbot is used to improve our products and services. This helps us provide more accurate information to customers and may be used to train Amazon AI models. This page describes how webmasters can control Amazonbot, Amzn-SearchBot, and Amzn-User interactions with their site. Each user agent setting is independent of the othersOne agent for product improvement and possible training. Amazon documents no finer split inside Amazonbot — the training opt-out is Amazonbot itself; search and user retrieval have their own agents. DOCUMENTED Amazon · Amazonbot
Common Crawl CCBot Open archive Common Crawl is a non-profit foundation founded with the goal of democratizing access to web information by producing and maintaining an open repository of web crawl data that is universally accessible and analyzable by anyone. Not a search engine and not an assistant. Blocking it affects an archive many parties use, model builders among them — not a product you can appear in. DOCUMENTED Common Crawl · CCBot

WHAT IS NOT DOCUMENTED

Two names that circulate, and no provider source for either

Two crawler names come up constantly and appear in no document by their provider. They are UNVERIFIED here, and they stay that way until the provider documents them. Third-party directories were found and refused as sources — several exist, and none is a provider source.

ProviderWhat we foundClassVerifiedSource
DeepSeek No crawler documentation by the provider was located. deepseek.com/robots.txt names no agent of its own. Third-party directories list a “DeepSeekBot”; none is a provider source. It stays unverified until DeepSeek documents it.
User-Agent: *
Allow: /
UNVERIFIED No provider source located.
ByteDance (“Bytespider”) No crawler documentation by the provider was located. The agent name circulates in third-party lists and server logs; neither is a provider statement. Same treatment. UNVERIFIED No provider source located.

Also not on this page

  • A recommended robots.txt. What to allow is a business decision about your content, not a technical default. This page gives you the classes so you can make it.
  • Any claim that allowing an agent produces a citation. Every provider that addresses the question declines to guarantee inclusion. Eligibility is not selection.
  • Crawler lists compiled by third parties. Several exist; none is a provider source, and this page has one sourcing rule: a DOCUMENTED row comes from the provider.
  • Server-log observations. We have not published crawler-hit logs. When we do, those rows will be OBSERVED and clearly separated from these.

LLMS.TXT

Google’s words, twice — and the file we maintain anyway

Google Search ignores it. In Google’s words:

You don’t need to create new machine readable files, AI text files, markup, or Markdown to appear in Google Search (including its generative AI capabilities), as Google Search itself doesn’t use them. Google · DOCUMENTED · 2026-09-05
It’s completely fine if you decide to create and maintain LLMS.txt files (or other similar files) for other services or systems that use these files. Doing so will neither harm nor help your site’s visibility or rankings in Google Search, as Google Search ignores them. Google · DOCUMENTED · 2026-09-05

That is the whole Google answer, and this page does not soften it. Which tools do read the file? Some do. This page does not list them, because we have not verified a single one from a provider source — any list here would be UNVERIFIED dressed as a reference. If we verify one, it gets a row, a link and a date like everything else.

We publish an llms.txt on this site and maintain it by hand, for the tools that read it. It is not part of any claim about Google.

Three neighbouring claims, same source, same date

Common claimGoogle’s wordsClassVerifiedSource
Content must be “chunked” into small pieces for AI. There’s no requirement to break your content into tiny pieces for AI to better understand it. DOCUMENTED Google · AI features and your website
Special structured data is needed for AI features. Structured data isn’t required for generative AI search, and there’s no special schema.org markup you need to add. DOCUMENTED Google · AI features and your website
A separate writing style is needed for AI search. The best practices for SEO continue to be relevant because our generative AI features on Google Search are rooted in our core Search ranking and quality systems. DOCUMENTED Google · AI features and your website

WHAT WE RUN ON OUR OWN SITE

Disclosure · OBSERVED · checkable against the live file

blackoarstudio.com/robots.txt · read

User-agent: *
Allow: /

Sitemap: https://blackoarstudio.com/sitemap-index.xml

Every agent on this page is currently allowed on this site, including the training bots. That is a deliberate position for a studio with no licensed archive to protect, and it is not advice. A publisher with content worth licensing might reasonably decide the opposite — the training table above is the table that lets them do so without losing search visibility, where the provider documents the separation.

OBSERVED — a live capture of our own file, 2026-09-05. Not a provider statement.

RE-VERIFICATION SCHEDULE

Next: 2026-12-05

This page describes other people’s systems and will go out of date without warning — in a way a broken-link check will not catch. Our own source list was compiled on 26 August 2026 and re-read on 2026-09-05. What the re-read found — OBSERVED:

  • Every URL on our own list of 26 August 2026 still resolved on 5 September 2026. A link check would have reported nothing.
  • In the same ten days, two agents appeared in provider documentation that were on no list of ours: OpenAI’s OAI-AdsBot and Meta’s Meta-WebIndexer. Neither moved a URL.
  • Google’s guide became explicit enough on llms.txt to quote verbatim. The URL did not change; the words did.
  • Two documentation homes moved (OpenAI’s bots page and Google’s crawler overview), both with a redirect. Both are recorded under their current address here.

Content drifts under stable URLs. That is why every row carries a date rather than just a link.

  • Every row carries the date its source was last read. That date is the claim.
  • Re-verified quarterly, and whenever a provider announces a change.
  • When a source moves, changes its wording or disappears, the row is marked and re-dated. The old claim is never silently edited — the same rule we apply to our own measurements.
  • Next scheduled re-verification: 2026-12-05.

SOURCES

Provider documents only · read 2026-09-05 · preserved with checksums

Two documents could not be captured raw and are marked above and in their rows. Both were read from the rendered page. Recorded so a reader knows which rows rest on a preserved document and which on a transcription. The IP endpoints in the verification table were fetched on the same date and preserved alongside the documents.

Reference version 1.0 · published 2026-09-05 · every source read 2026-09-05 · next re-verification 2026-12-05. A moved or changed source gets a marked, re-dated row; the old claim is never silently edited. The version number changes when a row’s substance changes, not when a sentence is reworded.