Skip to main content
Glama

batch_scrape

Scrape many URLs in a single parallel call with automatic deduplication and caching. Isolates per-URL failures so other pages still succeed.

Instructions

Scrape MANY URLs in ONE call (parallel, deduped, cache-aware).

Use when the user provides multiple URLs or you have a list of pages to fetch. More efficient than calling scrape N times.

Args: urls: Target URLs (deduped automatically; empties dropped). prefer: "auto" | "fast" | "stealth" | "llm". timeout: per-URL timeout in seconds. max_concurrency: parallel workers (default 4). include_html: include raw HTML per result (large; off by default).

Returns {requested, unique, succeeded, failed, results[]}. Per-URL failures are isolated — other URLs still succeed.

Returns: {requested, unique, succeeded, failed, results: [{url, markdown, ...}]}

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlsYes
preferNoauto
timeoutNo
include_htmlNo
max_concurrencyNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.7.1

TDQS

A4.7/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries full burden. It discloses key behaviors: parallel execution, automatic deduplication, cache-awareness, per-URL failure isolation, and the return structure. It does not mention rate limits, auth, or redirect handling, but the disclosed traits are sufficient for typical agent decisions. Minor gap is not explaining what 'cache-aware' implies operationally.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is efficiently written, front-loaded with the core purpose and usage guidance, then parameter details. It repeats the return structure twice (once in prose, once in a return block), which is slightly redundant but not harmful. Structure is logical and skimmable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool has 5 parameters and no annotations, the description covers all needed context: when to use, how to set parameters, return format, failure isolation, and efficiency rationale. It even explains the deduping and default for include_html. Nothing critical is missing for an agent to invoke it correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It does so thoroughly: each parameter gets a meaningful explanation (e.g., 'urls: Target URLs (deduped automatically; empties dropped)', 'prefer: "auto" | "fast" | "stealth" | "llm"', 'timeout: per-URL timeout in seconds'). This adds semantics well beyond the schema's types and defaults.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Scrape MANY URLs in ONE call' with explicit characteristics (parallel, deduped, cache-aware), clearly distinguishing it from the single-URL sibling 'scrape'. The verb-resource pairing is unmistakable and the emphasis on batch capability differentiates it from all sibling tools.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly states when to use it: 'Use when the user provides multiple URLs or you have a list of pages to fetch.' It also explains benefits over the alternative ('More efficient than calling `scrape` N times'), giving clear routing guidance without needing to inspect the schema.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.