Skip to main content
Glama

Scrape a list of URLs

scrape_urls
Idempotent

Fetch known URLs (1 to 500) once and return each page's content, markdown by default; no project is created. For one page with every format or browser steps use extract_url; to find pages by following links use crawl_site; to list a site's URLs without fetching them use map_site. Each page costs the credits of the engine that read it (1 plain fetch, 4 browser render), and a page the site refuses is free. A single URL answers in the same request; a list waits for the batch, and a large one comes back as an index with excerpts inside 60,000 characters. Hosted, a batch still going after the time budget comes back as a job for get_job, and repeating the same call returns that job instead of starting another.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlsYesThe pages to fetch: 1 to 500 absolute http(s) URLs.
configNoOptional fetch settings: any subset of the keys describe_project_config lists, e.g. {"render_js": "always", "only_main_content": true}. Omit for the defaults.
formatsNoComma-separated bodies to return per page, from markdown (default), text, cleanHtml and rawHtml, e.g. 'markdown,text'.markdown
parse_documentsNoRead PDFs, Word files and spreadsheets as text (default true); false leaves their text out. A parse_documents key in config wins.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed4 schema fields changed
    • addedInput schema / properties / config / description
      Added value: +"Optional fetch settings: any subset of the keys describe_project_config lists, e.g. {\"render_js\": \"always\", \"only_main_content\": true}. Omit for the defaults."
    • addedInput schema / properties / formats / description
      Added value: +"Comma-separated bodies to return per page, from markdown (default), text, cleanHtml and rawHtml, e.g. 'markdown,text'."
    • addedInput schema / properties / parse_documents / description
      Added value: +"Read PDFs, Word files and spreadsheets as text (default true); false leaves their text out. A parse_documents key in config wins."
    • addedInput schema / properties / urls / description
      Added value: +"The pages to fetch: 1 to 500 absolute http(s) URLs."
  2. First observed

TDQS

A4.6/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Goes well beyond the annotations by disclosing credit economics (1 plain fetch, 4 browser render), that refused pages are free, that a single URL answers in-request while a list waits for the batch, the 60,000-character excerpt limit, and that a hosted batch still running returns a job for get_job with repeat-call idempotency. This is rich operational context that the annotations alone do not convey.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Purpose and the sibling routing are front-loaded, and every sentence carries information about cost, batching or job behavior. It is dense and drifts into long run-on sentences at the end, but there is little pure filler to cut.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With no output schema, the description still explains the return shape (index with excerpts capped at 60,000 characters) and the asynchronous job path for large hosted batches. Combined with full parameter coverage, an agent has everything needed to call and handle this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents urls, config, formats and parse_documents in detail. The description repeats the 'markdown by default' default and echoes the batch/excerpt behavior but adds no new syntax or format guidance beyond the schema, making the baseline 3 appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource (fetch a list of known URLs, 1 to 500) plus the key scope fact that no project is created. It explicitly distinguishes itself from three named siblings (extract_url, crawl_site, map_site), so an agent can route correctly without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives explicit when-to-use-which routing: extract_url for one page with every format or browser steps, crawl_site for finding pages via links, map_site for listing URLs without fetching. Nothing about alternative selection is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.