| url | Yes | http(s) URL of the start page | |
| mode | No | The declared identity robots.txt, the page and the sitemaps are read under. authed is not offered: a map reads public sitemaps and one public page. | |
| debug | No | Return the full map response instead of the compact one. | |
| limit | No | Links returned at most. Default 5000; a hosted server takes up to 5000. Reaching it is status completed with stoppedBy limit. | |
| search | No | Keep only the URLs in which every word (at most 10) appears, case-insensitively, in the decoded URL or its title. A filter, not a ranking: the order stays the discovery order, and limit counts the matches. | |
| sitemap | No | include (default): the start page's links and the sitemaps. skip: no sitemap is read. only: no page is read; the links are the sitemap entries in their listed order (the start URL only when a sitemap lists it). | |
| timeout | No | Milliseconds for the whole map. Default 60000; a hosted server takes up to 60000. | |
| integration | No | Your own label for the integration or workflow this request belongs to (1 to 100 printable characters, no spaces). Stored in Octocrawl's records (the scrape record, the task status), never sent to the target. | |
| excludePaths | No | Pathname regexes that leave a URL out; they win over includePaths. | |
| includePaths | No | Pathname regexes a URL must match (as on crawl). | |
| regexOnFullURL | No | Match includePaths and excludePaths against the canonical URL instead of its pathname. Default false. | |
| ignoreRobotsTxt | No | Also return the URLs robots.txt disallows or whose robots.txt could not be read, each with that verdict (robots disallowed or unreachable), and read the start page and sitemaps past it. robots.txt is still read and recorded. Default false. A local server only; a hosted one refuses it. | |
| crawlEntireDomain | No | Admit URLs anywhere on the start host, not only in the start URL's path subtree. Default false. | |
| includeSubdomains | No | Admit every host under the start URL's apex (the host with one leading www. removed; no public-suffix list). Default false. Each new host's robots.txt is read, for at most 20 hosts. | |
| ignoreQueryParameters | No | Fold URLs that differ only in their query string into the first one seen, returned without its query; each merge is counted (refused.collapsed, with samples under debug). Default false. | |
| deduplicateSimilarURLs | No | Fold /a and /a/, / and /index.html, www and apex, http and https into one URL. Default true. | |