check_ai_crawlers
Check live which AI assistants may fetch a domain's pages from robots.txt, with per-assistant verdicts and the crawlers that decided each.
Instructions
Live check of which AI ASSISTANTS can fetch a site's pages, read from its robots.txt. assistantAccess is the answer — one verdict per assistant (ChatGPT, Claude, Perplexity, Microsoft Copilot, Google AI Overviews, Gemini Apps) with the crawlers that decided each named beside it. Count that array for totals; no count is stored, and there is deliberately no overall score. TWO THINGS IT DOES NOT TELL YOU, both of which get misreported: it says an assistant is PERMITTED to fetch the site, never that it cites it; and modelTrainingAccess is a separate, NEUTRAL fact — blocking training crawlers costs no visibility and is a legitimate content decision, so never report it as a gap or advise undoing it. The one exception is mechanical: where a token under modelTrainingAccess[].decidedByCrawlers also appears under assistantAccess[].decidedByCrawlers (Google-Extended is the documented case), that block DOES cost visibility — match on userAgentToken before applying the general rule. crawlers[].ruleAudience tells you whether a rule NAMED the crawler or a User-agent: * catch-all swept it up; the second is usually accidental and is the more actionable finding. When robots.txt cannot be read, it returns the read outcome and NO verdict — no assistant access, no crawler list, no advice — so check robotsTxt.read first. A failed read is not an open site.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| domain | Yes | Domain to scan, e.g. example.com | |
| industry | No | Industry context for benchmark comparison. |