Skip to main content
Glama

fetch_robots

Read-onlyIdempotent

Check if a target URL is crawlable for a specified user-agent by parsing robots.txt with Google's longest-match rule, then return the allow/deny verdict and matching rule.

Instructions

Fetch and parse the robots.txt for a given origin, then tell the caller whether a target URL is crawlable by a given user-agent. Follows the Google-style longest-match rule with Allow-wins-on-tie. Returns the raw robots.txt, parsed groups, sitemap references, and the allow/deny verdict with the matching rule.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesTarget URL to check. We derive origin + path automatically.
timeout_msNo
user_agentNoUser-agent string to match against groups (default '*')
max_redirectsNo
allow_private_hostsNoAllow loopback / private / link-local targets for this call (default false). Refused unless the server operator launched fetch-mcp with FETCH_MCP_ALLOW_PRIVATE_HOSTS=1 -- SSRF protection stays on by default either way.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv0.8.0
    • addedInput schema / properties / allow_private_hosts / description
      Added value: +"Allow loopback / private / link-local targets for this call (default false). Refused unless the server operator launched fetch-mcp with FETCH_MCP_ALLOW_PRIVATE_HOSTS=1 -- SSRF protection stays on by default either way."
  2. First observedv0.6.1

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Beyond the annotations (readOnly, openWorld, idempotent, non-destructive), the description adds the specific parsing rule (longest-match with Allow-wins-on-tie) and the returned data (raw robots.txt, parsed groups, sitemap references, and verdict). It does not conflict with annotations and adds meaningful behavioral context.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is two sentences, front-loaded with the core purpose, then the matching rule, then the return contents. Every sentence adds necessary information with no filler. It is well-structured and appropriately sized.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given there is no output schema, the description appropriately covers the return values (raw robots.txt, parsed groups, sitemap references, verdict). It also explains the matching logic. It does not mention error handling or network behavior, but these are likely not critical for a read-only tool with annotations indicating no side effects. The description is sufficient for an agent to know what the tool does and what it returns.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema already describes url, user_agent, and allow_private_hosts. The description only indirectly references url and user_agent without adding new meaning. The timeout_ms and max_redirects parameters lack descriptions in both schema and description, and the description does not compensate for this gap. With 60% schema coverage, the baseline is 3, and the description adds little parameter-specific value.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states it fetches and parses robots.txt and returns a crawlability verdict for a URL and user-agent. It uses specific verbs and resources, and the function is distinct from sibling fetch tools by focusing on robots.txt, not generic content or sitemaps.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the primary use case: checking crawlability for a given URL and user-agent. It provides clear context for when to use this tool, though it does not explicitly name alternatives or provide exclusions. Since no sibling handles robots.txt specifically, the implied usage is sufficient.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.