Skip to main content
Glama
Akxan
by Akxan

AI crawler access check

ai_crawler_access
Read-onlyIdempotent

Check if AI and search crawlers can access any page by testing robots.txt and live bot requests. Detect UA-based blocks that prevent AI assistants from citing your site.

Instructions

Check whether AI and search crawlers (GPTBot, OAI-SearchBot, ChatGPT-User, ClaudeBot, Claude-SearchBot, PerplexityBot, Google-Extended, Googlebot, Bingbot, Applebot, Amazonbot, Meta, CCBot, Bytespider...) can reach a page: robots.txt rules for the site root and for the URL, plus a live request with each bot's User-Agent to detect UA-based blocks (e.g. Cloudflare 'block AI bots', WAF rules). Being blocked from OAI-SearchBot, PerplexityBot or Claude-SearchBot means the site cannot be cited by those assistants.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesA representative page, e.g. the homepage or an important article.
liveFetchNoAlso request the page with each bot UA (one request per bot).

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.5.1
    • removedInput schema / additionalProperties
      Removed value: -false
  2. First observedv0.3.0

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The annotations already establish that the tool is read-only, idempotent, and non-destructive. The description adds meaningful behavioral detail beyond those annotations: it performs live HTTP requests with each bot's User-Agent, checks robots.txt at both root and URL levels, and can detect UA-based blocks like Cloudflare or WAF rules. This is valuable context that annotations alone do not provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two dense sentences with no filler. The bot list and the citation-blocking consequence are informative rather than redundant, and the primary action is front-loaded. Every phrase adds useful detail for tool selection and invocation.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given that there is no output schema, the description does a good job of conveying what the tool checks and why, including the robots.txt and live-request dimensions. However, it does not explicitly describe the return format or note that liveFetch=true may generate many per-bot requests, so a small gap remains.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already covers both parameters at 100%, including the meaning of 'url' and 'liveFetch' with 'one request per bot'. The description reinforces these semantics but does not materially expand on them beyond naming specific bot UAs, so the baseline of 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb and resource ('Check whether AI and search crawlers can reach a page'), enumerates the exact crawlers involved, and explains the two concrete mechanisms (robots.txt and live UA-based requests). It is clearly differentiated from generic robots checks and citation tools by naming the crawler-access scope and the consequence for assistant citation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides clear context by stating that being blocked from certain bots means assistants cannot cite the site, which tells the agent when this tool is relevant. It does not explicitly name sibling alternatives like robots_check or ai_citation_check or give exclusions, but the intended use case is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.