Skip to main content
Glama

web_crawl

Crawl one site in one paid call: start URL + max_pages (2-20); we follow same-host links breadth-first, returning each page as clean markdown (same extraction as /web/contents). Priced at $0.002 per requested page, quoted as max_pages up front — finding fewer pages is still a complete delivery. Failed pages are skipped and listed in skipped[]; the call still settles. Only an unreadable start page is a 400, no charge. USDC on Base, no account, no API key.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesThe http(s) start URL of the site to crawl. Literal IP addresses, localhost and non-http(s) schemes are rejected.
max_pagesNoMaximum pages to crawl, breadth-first over same-host links (2-20, default 5). Priced at $0.001 per requested page; finding fewer pages is still a complete delivery.
max_chars_per_pageNoMaximum content length per page in characters (1000-20000, default 5000); longer pages are cut and flagged truncated

Schema Changelog

Changes observed during successful MCP inspections.

  1. Added

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and discloses a lot: per-page pricing, up-front quoting, settlement despite skipped pages, the 400-on-unreadable-start error behavior, and the USDC-on-Base no-account/no-API-key model. However, it quotes $0.002 per page while the max_pages schema description says $0.001 — a 2x pricing inconsistency that leaves an agent with conflicting cost data.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Five dense sentences, each earning its place: core behavior, pricing model, failure handling, error semantics, and payment/auth. The core purpose is front-loaded and there is no filler or redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a tool with no annotations and no output schema, the description covers the essentials: output format, skipped-page handling, error conditions, and the cost model. The main gap is the response envelope — pages and skipped[] are mentioned, but the overall return container is never described.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3; the schema already documents types, ranges, defaults, and the truncation flag for all three parameters. The description adds billing semantics (quoted at max_pages up front, fewer pages still a complete delivery) but repeats the price with a conflicting figure and contributes nothing on max_chars_per_page.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific action ('Crawl one site in one paid call') with a clear resource (one site), traversal rule (breadth-first over same-host links), and output format (clean markdown per page). The 'same extraction as /web/contents' reference situates it against the sibling web tools without confusing it with single-page fetchers.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The scope — 'one site in one paid call' with same-host breadth-first traversal — gives an agent clear context for when this tool fits: multi-page extraction within a single domain. It does not explicitly name alternatives (e.g., web_contents for a single page) or state when not to use it, so it stops short of full routing guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources