GPTBot (OpenAI), Google-Extended (Gemini), Anthropic-AI (Claude), and CCBot (Common Crawl) each have dedicated user-agent strings. You can selectively block AI training crawlers while keeping search bots allowed — granular control matters.
If your content is already indexed by search engines, AI answer engines can still cite it via search APIs (like Google's Grounding API). Robots.txt controls crawling, not citation. To actually control AI usage, combine robots.txt with meta robots and data-noai headers.
Google ignores Crawl-delay (use Search Console instead), but Bing, Yandex, and most AI crawlers honor it. If your server is under load from AI bots, adding a crawl-delay for their user-agents is a practical throttle.