Skip to main content
Glama

AIsa Web Search & Research

Submit an asynchronous batch scrape job.

post_firecrawl_batch_scrape
Read-onlyIdempotent

Scrape many URLs as one background job. urls and an Idempotency-Key are required; maxConcurrency, onlyMainContent, includeTags, excludeTags, maxAge, minAge and timeout tune it. Asynchronous. Submitting returns HTTP 202 and a job envelope — id, object, endpoint, status, createdAt, completedAt, pricing, output, error — with output still null. Poll get_firecrawl_batch_scrape_job until terminal; output is then an array of documents with markdown and metadata. Use it when you have a list of URLs and do not need them immediately. When you do need them immediately, post_tavily_extract returns a small batch synchronously in about a second; for a single page post_firecrawl_scrape. Send a fresh Idempotency-Key per distinct batch.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlsYes1 to 1000 unique HTTPS URLs to scrape. PDF URLs are not supported.
maxAgeNoMaximum acceptable cache age in milliseconds.
minAgeNoMinimum cache age in milliseconds before a page is refetched.
timeoutNoPer-page timeout in milliseconds.
excludeTagsNoHTML tags/selectors to drop.
includeTagsNoHTML tags/selectors to keep.
maxConcurrencyNoMaximum number of concurrent scrapes (1 to 20).
Idempotency-KeyYesUnique key (1 to 191 characters) that makes the submit idempotent. Re-submitting with the same key returns the original job.
onlyMainContentNoReturn only the main content of each page.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.4/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint=true, openWorldHint=true, idempotentHint=true, and destructiveHint=false, so the safety profile is covered. The description adds valuable behavioral context beyond that: the HTTP 202 submission semantics, the job envelope shape with output initially null, the terminal-state output being an array of markdown/metadata documents, and the idempotency behavior of re-submitting with the same key returning the original job. It does not contradict the annotations. It loses one point because it doesn't mention polling frequency or error/retry behavior, but the core async lifecycle is well disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is dense but well-organized: it front-loads the core action and required fields, then the async behavior, then the polling instruction, then the when-to-use alternatives. Every sentence earns its place. It is slightly long, but the length is justified by the async lifecycle and routing guidance. It loses one point because the job-envelope field list is somewhat verbose and could be trimmed without losing meaning.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 9-parameter async tool with a rich output schema, the description is complete: it covers required fields, async submission semantics, the polling endpoint, the terminal output shape, and the alternative tools. The output schema exists, so the description needn't explain return values in detail. An agent has everything needed to select and invoke this tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all 9 parameters. The description adds a little by naming the tuning parameters and emphasizing that urls and Idempotency-Key are required, but it doesn't add meaning beyond the schema for individual parameters. Baseline 3 is appropriate when the schema carries the full parameter documentation burden.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb and resource: 'Scrape many URLs as one background job.' It clearly distinguishes this from siblings by naming post_tavily_extract and post_firecrawl_scrape as alternatives for immediate or single-page needs. The async nature and job-envelope return are stated up front, so an agent can tell this tool apart from the synchronous scraping siblings without opening the schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description gives explicit when-to-use guidance: 'Use it when you have a list of URLs and do not need them immediately.' It also names the alternatives and the conditions that select them: post_tavily_extract for a small synchronous batch, post_firecrawl_scrape for a single page. It even instructs to poll get_firecrawl_batch_scrape_job until terminal, which is the correct follow-up action.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources