Skip to main content
Glama
crawlbrulee

@crawlbrulee/mcp

Official
by crawlbrulee

Map a website

map

Enumerate a website's URLs before scraping by combining sitemap discovery with homepage link extraction. Get paginated, filterable lists of internal, external, and subdomain links.

Instructions

Build (or fetch a cached) link-map for a website by combining sitemap discovery with homepage link extraction. Returns paginated lists of discovered URLs, each item just { url }. Filterable by link type (internal / external / subdomain). Use this to enumerate a site before scraping selected pages. max_urls (default 5000, max 100000) is a discovery budget, not a trim at the end: discovery stops as soon as that many URLs are found, so a smaller value is a faster, cheaper crawl. A map stopped that way returns exactly max_urls links with response_capped false — the signal that the site has more is response_meta.truncation.discovery_cap_reason. When that is "max_urls", ask again with a higher max_urls to get more; "unread_files" means a sitemap file could not be read this time and is often temporary, so asking again later can return more; "time", "file_budget", "depth" and "file_size" mean the site itself is big, slow or deep and a retry will not help. discovery_capped says discovery stopped early, sitemaps_skipped how many sitemap files were skipped or only partly read. limit (default 5000, max 10000) only pages the answer. Returned URLs are normalized the same way scrape normalizes its returned url, so map-then-scrape stays on one host. Results are ordered with the most useful links first. The response carries response_meta.usage = { credits, engine, proxy } — the resolved proxy tier is never auto; map responses do not include screenshot-slice accounting.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesThe website URL to map. Mapping always targets the site root, so the path, query string and fragment are dropped; known tracking parameters are removed before the request is processed.
pageNoPage number for paginated results
cacheNoCache settings for this request
limitNoNumber of URLs to return per page. Default 5000, maximum 10000.
proxyNoProxy tier to use for fetchingauto
typesNoFilter which link types to include
locationNoOptional country emulation for the map
max_urlsNoMaximum number of URLs to discover and store in the map. Default 5000, maximum 100000. Sitemap discovery stops as soon as this many URLs have been found, so a smaller value is a faster and lighter crawl, not just a smaller answer.
sitemap_onlyNoOnly use sitemap.xml — skip homepage link extraction

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
linksYesList of discovered URLs for the current page
response_metaYesResponse metadata including pagination, truncation, and usage info

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv1.0.3

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so thoroughly. It discloses discovery-budget stopping semantics, `response_capped` false behavior, all truncation reasons, URL normalization relative to `scrape`, ordering by usefulness, and response details like `response_meta.usage` and that `proxy` is never `auto`. This goes well beyond the schema and gives agents a reliable mental model of non-obvious behavior.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but every sentence earns its place. It is front-loaded with the core function and use case, then progresses to parameters, then to response/error nuances. No filler or redundancy; the density is justified by the tool's complexity and the absence of annotations.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool with 9 parameters, nested objects, and an output schema, the description is complete. It covers the return format, pagination, cache semantics, filtering, discovery stopping, truncation signals, retry guidance, and response metadata. An agent has everything needed to decide when to call it and how to interpret results.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds meaningful nuance beyond the schema, particularly for `max_urls` (a discovery budget, not a trim), `limit` (only pages the answer), and the behavior of `types` filtering. It does not add supplementary meaning for every parameter, but the extra context is valuable and improves correct usage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Build (or fetch a cached) link-map for a website' via sitemap discovery and homepage link extraction. It explicitly states what the tool returns (paginated lists of discovered URLs) and positions it as an enumeration tool distinct from the scraping siblings.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly says 'Use this to enumerate a site before scraping selected pages', giving a clear context. It also explains when retrying is useful based on `discovery_cap_reason` (e.g., 'max_urls' warrants a higher budget, 'unread_files' may be temporary). It does not explicitly name an alternative tool or provide a when-not-to-use statement, but the use case is clear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.