Skip to main content
Glama

Crawl a whole site once

crawl_site
Idempotent

Crawl a site once from a start URL -- following its links and reading its sitemap, up to limit pages -- and return what it found, without setting up a project. Use it to read a site or section whose page URLs you do not know; for known URLs use scrape_urls, and to watch a site over time use create_project (or keep_crawl_as_project on this crawl afterwards). Each page costs the credits of the engine that read it (usually 1 to 4) and refused pages are free. It waits for the crawl, minutes for a large limit; hosted, a crawl still going after the time budget comes back as a job for get_job, and repeating the call returns the same crawl. The answer is an index of the pages read plus excerpts inside 60,000 characters; get_job with url reads one page in full and with cursor the next window. The crawl is kept for a day; pass the result's crawl_id to keep_crawl_as_project to keep it for good.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesThe absolute http(s) URL to start from, e.g. the home page.
limitNoThe most pages to read, 1 to 5,000 (default 50); the plan's page cap also applies.
configNoOptional fetch settings: any subset of the keys describe_project_config lists, e.g. {"render_js": "always", "only_main_content": true}. Omit for the defaults.
max_depthNoHow many links deep to follow from the start URL (0 = that page only; default 3).
exclude_pathsNoGlobs over the URL path to skip, e.g. ['/tag/*', '*.pdf']; an exclude wins over an include.
include_pathsNoGlobs over the URL path to keep, e.g. ['/blog/*'] for a section (its index included); a bare '/blog/' matches only that one page. Omit for the whole site.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed6 schema fields changed
    • addedInput schema / properties / config / description
      Added value: +"Optional fetch settings: any subset of the keys describe_project_config lists, e.g. {\"render_js\": \"always\", \"only_main_content\": true}. Omit for the defaults."
    • addedInput schema / properties / exclude_paths / description
      Added value: +"Globs over the URL path to skip, e.g. ['/tag/*', '*.pdf']; an exclude wins over an include."
    • addedInput schema / properties / include_paths / description
      Added value: +"Globs over the URL path to keep, e.g. ['/blog/*'] for a section (its index included); a bare '/blog/' matches only that one page. Omit for the whole site."
    • addedInput schema / properties / limit / description
      Added value: +"The most pages to read, 1 to 5,000 (default 50); the plan's page cap also applies."
    • addedInput schema / properties / max_depth / description
      Added value: +"How many links deep to follow from the start URL (0 = that page only; default 3)."
    • addedInput schema / properties / url / description
      Added value: +"The absolute http(s) URL to start from, e.g. the home page."
  2. First observed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations cover safety (idempotent, non-destructive) but the description adds substantial behavioral context beyond them: per-page credit costs (1-4, refused pages free), synchronous waiting with a time budget, hosted crawls returning a job retrievable via get_job, idempotent repeat-call behavior, and a one-day retention window. This is far past what annotations alone convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and alternatives are front-loaded and every sentence carries information. It is, however, a dense run-on block with several clauses chained by semicolons, which slightly hurts scannability for an agent parsing it quickly.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Despite no output schema, the description explains the return shape (index of pages plus excerpts capped at 60,000 characters), how to page through one page via get_job with url vs cursor, and the crawl_id retention handoff. An agent has enough to call and consume the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents url, limit, config, max_depth, and include/exclude paths thoroughly. The description reinforces `limit` and the crawl_id return but adds no parameter syntax or format detail beyond the schema. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (crawl) and resource (site from a start URL), including the mechanism (following links, reading sitemap, up to `limit` pages) and the key scoping outcome (returns an index without creating a project). It explicitly distinguishes itself from scrape_urls and create_project. An agent can select it correctly without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use ('whose page URLs you do not know') and when-not, naming the alternatives: scrape_urls for known URLs, create_project to watch a site over time, keep_crawl_as_project to persist this crawl. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.