Skip to main content
Glama

octocrawl

map

Read-onlyIdempotent

List a site's URLs without fetching each page: the start URL, the links on its page (read on the http lane alone; no browser) and the entries of the sitemaps the site declares (robots.txt Sitemap: lines, else /sitemap.xml), inside one deadline. Every URL is in the crawl's scope (the start host and its www twin, the start URL's path subtree, assets left out, similar URLs folded) and allowed by its host's robots.txt unless ignoreRobotsTxt is set; what was left out is counted. A title is never fetched: the start page's own, an anchor's text or a sitemap's news title. At the deadline the answer is what was found, status partial (failed when nothing), stoppedBy timeout. Compact by default ({ id, status, stoppedBy, links: [{ url, title?, description?, robots? }], warning?, agentHints?, counts }; robots only on a link robots.txt keeps out, under ignoreRobotsTxt); debug=true returns the full map with each link's evidence (via, sitemapFile, lastmod, robots), the sources read and the refusals. One page body is read at most: a site without a sitemap maps only its start page's links; crawl reads further pages.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYeshttp(s) URL of the start page
modeNoThe declared identity robots.txt, the page and the sitemaps are read under. authed is not offered: a map reads public sitemaps and one public page.
debugNoReturn the full map response instead of the compact one.
limitNoLinks returned at most. Default 5000; a hosted server takes up to 5000. Reaching it is status completed with stoppedBy limit.
searchNoKeep only the URLs in which every word (at most 10) appears, case-insensitively, in the decoded URL or its title. A filter, not a ranking: the order stays the discovery order, and limit counts the matches.
sitemapNoinclude (default): the start page's links and the sitemaps. skip: no sitemap is read. only: no page is read; the links are the sitemap entries in their listed order (the start URL only when a sitemap lists it).
timeoutNoMilliseconds for the whole map. Default 60000; a hosted server takes up to 60000.
integrationNoYour own label for the integration or workflow this request belongs to (1 to 100 printable characters, no spaces). Stored in Octocrawl's records (the scrape record, the task status), never sent to the target.
excludePathsNoPathname regexes that leave a URL out; they win over includePaths.
includePathsNoPathname regexes a URL must match (as on crawl).
regexOnFullURLNoMatch includePaths and excludePaths against the canonical URL instead of its pathname. Default false.
ignoreRobotsTxtNoAlso return the URLs robots.txt disallows or whose robots.txt could not be read, each with that verdict (robots disallowed or unreachable), and read the start page and sitemaps past it. robots.txt is still read and recorded. Default false. A local server only; a hosted one refuses it.
crawlEntireDomainNoAdmit URLs anywhere on the start host, not only in the start URL's path subtree. Default false.
includeSubdomainsNoAdmit every host under the start URL's apex (the host with one leading www. removed; no public-suffix list). Default false. Each new host's robots.txt is read, for at most 20 hosts.
ignoreQueryParametersNoFold URLs that differ only in their query string into the first one seen, returned without its query; each merge is counted (refused.collapsed, with samples under debug). Default false.
deduplicateSimilarURLsNoFold /a and /a/, / and /index.html, www and apex, http and https into one URL. Default true.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
idYesThe map's record id (GET /v1/maps/:id on the REST API).
linksYes
countsNo
statusYes
warningNoThe warnings' messages, joined.
stoppedByYes
agentHintsNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

Score is being calculated.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.