Skip to main content
Glama

site_extract

Render a whole section of a site - up to 25 pages - in one call and get every page back as clean text, with JavaScript executed. Same host only; robots.txt is obeyed and anything it disallows is skipped and named. FREE: the crawl PLAN - exactly which pages would be fetched and what robots.txt allows - so you can see what you would get before paying. PAID ($0.14, x402 on Arc/Base/Solana/X1): every page rendered and returned. It costs more than one page because it is up to 25 real browser renders, not a lookup. If a bot wall on the site lets the first page through but blocks the rest, the call fails with poor-yield and names every page it could not render, instead of billing you for a list of failures.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesthe page to start from
fullNotrue asks for the paid crawl. Without a payment you get a 402 carrying the price and every rail we accept; sign it and call again with `payment`.
pagesNohow many pages, default 10, hard cap 25
paymentNoa signed x402 payment (the same base64 payload you would put in the PAYMENT-SIGNATURE header). Pass it here and the purchase completes inside this tool call.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed2 schema fields changed
    • changedInput schema / properties / full / description
      Previous value: -"true asks for the paid crawl; you will get a 402 with payment options"New value: +"true asks for the paid crawl. Without a payment you get a 402 carrying the price and every rail we accept; sign it and call again with `payment`."
    • addedInput schema / properties / payment
      Added value: +{
      +  "description": "a signed x402 payment (the same base64 payload you would put in the PAYMENT-SIGNATURE header). Pass it here and the purchase completes inside this tool call.",
      +  "type": "string"
      +}
  2. Added

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It discloses several critical behaviors: respects robots.txt and skips disallowed pages, costs more than one page due to real browser renders, and has a failure mode (poor-yield) that returns page names instead of billing. However, it does not mention rate limits, timeouts, or whether links are followed recursively and how depth is determined, leaving room for ambiguity.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense with information but each sentence earns its place: scope, host restriction, robots.txt handling, free plan, paid plan, cost justification, and failure mode. It is front-loaded with the main purpose and then details. Slightly long but justified given the complexity.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 4 parameters (one required), no output schema, and no annotations, the description covers most essentials: what it does, how payment works, error handling, and constraints. Missing: what the output format looks like (plain text), how it handles multi-page navigation (e.g., link discovery), and whether authentication/cookies are supported. But for an agent to call it, the description is fairly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so baseline is 3. The description adds valuable meaning: it explains 'full' triggers the paid crawl, and the free version returns a crawl plan, and 'payment' is a signed x402 payload. It also clarifies 'pages' has a hard cap of 25 and default 10. This goes beyond the schema's terse field names.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool renders a whole section of a site (up to 25 pages) in one call, returning clean text with JavaScript executed. It explicitly contrasts with the sibling page_extract by emphasizing the whole-section scope and the multiple-page coverage, making the distinction obvious.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly explains when to use this tool (when you need multiple pages from the same host) and notes the paid vs free flow: the free crawl plan shows what would be fetched, and the paid version returns all pages. It also warns about bot walls and the poor-yield failure mode, making the usage context explicit.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources