fetch_robots
Check if a target URL is crawlable for a specified user-agent by parsing robots.txt with Google's longest-match rule, then return the allow/deny verdict and matching rule.
Instructions
Fetch and parse the robots.txt for a given origin, then tell the caller whether a target URL is crawlable by a given user-agent. Follows the Google-style longest-match rule with Allow-wins-on-tie. Returns the raw robots.txt, parsed groups, sitemap references, and the allow/deny verdict with the matching rule.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Target URL to check. We derive origin + path automatically. | |
| timeout_ms | No | ||
| user_agent | No | User-agent string to match against groups (default '*') | |
| max_redirects | No | ||
| allow_private_hosts | No | Allow loopback / private / link-local targets for this call (default false). Refused unless the server operator launched fetch-mcp with FETCH_MCP_ALLOW_PRIVATE_HOSTS=1 -- SSRF protection stays on by default either way. |