GEOKit
🕷️

AI Crawlers Directory

24 Crawlers

Complete directory of all known AI crawlers with User-Agent strings, IP ranges, and ready-to-use robots.txt snippets.

OpenAI / GPTBot

training

Web crawler for training OpenAI's models.

User-Agent:GPTBot

OpenAI / ChatGPT-User

browsing

Used by ChatGPT plugins and web browsing feature.

User-Agent:ChatGPT-User

OpenAI / OAI-SearchBot

browsing

Crawler for OpenAI Search (SearchGPT) functions.

User-Agent:OAI-SearchBot

Anthropic / ClaudeBot

training

Used by Anthropic to crawl the web for training Claude models.

User-Agent:ClaudeBot

Anthropic / anthropic-ai

training

Legacy/alternate user agent for Anthropic.

User-Agent:anthropic-ai
Opt-out: Yes

Anthropic / Claude-Web

browsing

Used by Claude to browse live web pages for users.

User-Agent:Claude-Web
Opt-out: Yes

Google / Google-Extended

training

Controls whether your site helps improve Google's AI models (Gemini).

User-Agent:Google-Extended

Google / Googlebot

indexing

Primary crawler for Google Search (indexing). Also used to surface in AI Overviews.

User-Agent:Googlebot

Microsoft / Bingbot

indexing

Primary crawler for Bing Search and Copilot features.

User-Agent:bingbot

Microsoft / MSNBot

indexing

Legacy Microsoft crawler.

User-Agent:msnbot
Opt-out: Yes

Perplexity / PerplexityBot

browsing

Web crawler for Perplexity AI search engine.

User-Agent:PerplexityBot

Meta / Meta-ExternalAgent

training

Crawls web pages linked within Meta services for AI features.

User-Agent:Meta-ExternalAgent

Meta / Meta-ExternalFetcher

browsing

Fetches preview links on Meta platforms.

User-Agent:Meta-ExternalFetcher
Opt-out: Yes

Meta / FacebookBot

browsing

Used to fetch link previews for sharing on Facebook/Messenger.

User-Agent:facebookexternalhit
Opt-out: Yes

Apple / Applebot

indexing

Used by Siri and Spotlight suggestions.

User-Agent:Applebot

Apple / Applebot-Extended

training

Opt-out specific crawler for Apple's foundation models.

User-Agent:Applebot-Extended

ByteDance / Bytespider

both

ByteDance's generic crawler for various purposes including AI.

User-Agent:Bytespider
Opt-out: Yes

You.com / YouBot

browsing

Crawler for the You.com AI search engine.

User-Agent:YouBot

DuckDuckGo / DuckAssistBot

browsing

Used by DuckDuckGo's AI-assisted answers.

User-Agent:DuckAssistBot

Cohere / Cohere-ai

training

Data collection crawler for Cohere's language models.

User-Agent:cohere-ai

Diffbot / Diffbot

indexing

Knowledge Graph and web data extraction API.

User-Agent:Diffbot

Common Crawl / CCBot

training

Builds an open repository of web crawl data, heavily used for training AI models like GPT and Llama.

User-Agent:CCBot

Amazon / Amazonbot

both

Amazon's web crawler used for indexing and potentially AI training.

User-Agent:Amazonbot

Oracle / Grapeshot

indexing

Contextual targeting crawler for advertising.

User-Agent:grapeshot
Opt-out: YesOpt-Out Page ↗

Frequently Asked Questions

What is an AI crawler?

AI crawlers (or bots) are automated programs used by AI companies to scrape web pages. They generally serve two purposes: training (gathering massive amounts of text to train models like GPT-4) and browsing (fetching live information when a user asks a real-time question).

Will blocking AI crawlers hurt my SEO?

It depends on the bot. Blocking training-only bots (like GPTBot or CCBot) won't impact traditional search rankings, but you won't be included in their base models. Blocking indexing bots (like Googlebot or Bingbot) will remove your site from search results. Be very careful with Search Engine crawlers.

How do I allow only specific AI crawlers?

You can use the generator on the right to select the crawlers you want to interact with. A common strategy is to allow browsing bots (so AI can cite your site in real-time answers) but block training bots (if you want to protect your intellectual property).

robots.txt Generator

Action for 24 selected bots:

24 of 24 selected
# Generated by GEOKit AI Crawlers Directory

User-agent: GPTBot
Disallow: /

User-agent: ChatGPT-User
Disallow: /

User-agent: OAI-SearchBot
Disallow: /

User-agent: ClaudeBot
Disallow: /

User-agent: anthropic-ai
Disallow: /

User-agent: Claude-Web
Disallow: /

User-agent: Google-Extended
Disallow: /

User-agent: Googlebot
Disallow: /

User-agent: bingbot
Disallow: /

User-agent: msnbot
Disallow: /

User-agent: PerplexityBot
Disallow: /

User-agent: Meta-ExternalAgent
Disallow: /

User-agent: Meta-ExternalFetcher
Disallow: /

User-agent: facebookexternalhit
Disallow: /

User-agent: Applebot
Disallow: /

User-agent: Applebot-Extended
Disallow: /

User-agent: Bytespider
Disallow: /

User-agent: YouBot
Disallow: /

User-agent: DuckAssistBot
Disallow: /

User-agent: cohere-ai
Disallow: /

User-agent: Diffbot
Disallow: /

User-agent: CCBot
Disallow: /

User-agent: Amazonbot
Disallow: /

User-agent: grapeshot
Disallow: /

Copy this snippet into your website's robots.txt file to apply the rules.

Share this tool

Found this tool useful? Help others by linking to it from your blog or resources page.