Map a website
mapEnumerate a website's URLs before scraping by combining sitemap discovery with homepage link extraction. Get paginated, filterable lists of internal, external, and subdomain links.
Instructions
Build (or fetch a cached) link-map for a website by combining sitemap discovery with homepage link extraction. Returns paginated lists of discovered URLs, each item just { url }. Filterable by link type (internal / external / subdomain). Use this to enumerate a site before scraping selected pages. max_urls (default 5000, max 100000) is a discovery budget, not a trim at the end: discovery stops as soon as that many URLs are found, so a smaller value is a faster, cheaper crawl. A map stopped that way returns exactly max_urls links with response_capped false — the signal that the site has more is response_meta.truncation.discovery_cap_reason. When that is "max_urls", ask again with a higher max_urls to get more; "unread_files" means a sitemap file could not be read this time and is often temporary, so asking again later can return more; "time", "file_budget", "depth" and "file_size" mean the site itself is big, slow or deep and a retry will not help. discovery_capped says discovery stopped early, sitemaps_skipped how many sitemap files were skipped or only partly read. limit (default 5000, max 10000) only pages the answer. Returned URLs are normalized the same way scrape normalizes its returned url, so map-then-scrape stays on one host. Results are ordered with the most useful links first. The response carries response_meta.usage = { credits, engine, proxy } — the resolved proxy tier is never auto; map responses do not include screenshot-slice accounting.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The website URL to map. Mapping always targets the site root, so the path, query string and fragment are dropped; known tracking parameters are removed before the request is processed. | |
| page | No | Page number for paginated results | |
| cache | No | Cache settings for this request | |
| limit | No | Number of URLs to return per page. Default 5000, maximum 10000. | |
| proxy | No | Proxy tier to use for fetching | auto |
| types | No | Filter which link types to include | |
| location | No | Optional country emulation for the map | |
| max_urls | No | Maximum number of URLs to discover and store in the map. Default 5000, maximum 100000. Sitemap discovery stops as soon as this many URLs have been found, so a smaller value is a faster and lighter crawl, not just a smaller answer. | |
| sitemap_only | No | Only use sitemap.xml — skip homepage link extraction |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| links | Yes | List of discovered URLs for the current page | |
| response_meta | Yes | Response metadata including pagination, truncation, and usage info |