Skip to main content
Glama

Site Check

Check AI crawler access

check_ai_crawler_access
Read-onlyIdempotent

Checks whether 14 AI crawlers may fetch a page. Use it when the user asks whether AI crawlers can read a website or page, for example "can AI crawlers read example.com?", "is GPTBot blocked in my robots.txt?" or "which AI bots are allowed on https://example.com/blog?". Pass the page address. Returns, for 14 AI crawlers, whether robots.txt allows the page and the rule that decided it, plus the sitemaps, the redirect chain, and the page's X-Robots-Tag header and robots meta tag. It reads robots.txt only and does not test whether a firewall blocks crawlers. Do not use it for private or local addresses, for security scans, or to collect a site's content.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesAddress of a public web page, for example https://example.com/pricing. A bare domain such as example.com means https. Only public http and https sites on the standard ports are checked.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYes
notesNo
noticeYes
robotsNo
sourceYes
statusNo
failureNo
crawlersNo
finalUrlNo
reachableYes
redirectsNo
metaRobotsNo
xRobotsTagNo
blockedCrawlersNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover readOnly, idempotent, openWorld and non-destructive, but the description adds material behavior beyond them: it reads robots.txt only, it does not test whether a firewall blocks crawlers, and it enumerates exactly what it inspects (robots.txt rules, sitemaps, redirect chain, X-Robots-Tag, robots meta). Those caveats are what an agent needs to avoid over-trusting the result.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the purpose and use cases, then limitations; every sentence carries information. The enumeration of return fields is somewhat verbose given a full output schema exists, which slightly blunts conciseness but does not bury the key guidance.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a single-parameter, read-only probe with annotations and an output schema already supplied, the description covers triggers, exclusions, method limits, and result contents. Nothing an agent needs in order to decide to call it or interpret it is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and the single parameter already documents the URL format, bare-domain-to-https behavior, and public/standard-port restriction. The description's 'Pass the page address' and the query examples add no syntax or format detail beyond the schema, so the baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Checks whether 14 AI crawlers may fetch a page') and scopes it to robots.txt-based access, which cleanly separates it from siblings like ai_readiness_check, check_page_tags, and trace_redirects. The named crawler count and the page-level granularity make the purpose unambiguous.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives concrete trigger phrasing ('can AI crawlers read example.com?', 'is GPTBot blocked in my robots.txt?') plus explicit exclusions: no private/local addresses, no security scans, no content collection. It also names the boundary condition of the method (robots.txt only, no firewall testing), so an agent knows when this tool is the wrong choice.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources