Skip to main content
Glama

AIsa Web Search & Research

Submit an asynchronous crawl job.

post_firecrawl_crawl
Read-onlyIdempotent

Crawl a whole site rooted at url and return the content of every page it keeps. url, limit and an Idempotency-Key are required; steer it with includePaths, excludePaths, maxDiscoveryDepth, crawlEntireDomain, allowSubdomains, delay and maxConcurrency. Asynchronous. Submitting returns HTTP 202 and a job envelope — id, object, endpoint, status, createdAt, completedAt, pricing, output, error — with output still null. Poll get_firecrawl_crawl_job until status is completed, failed or cancelled; output is then an array of pages, each with markdown and metadata. A 3-page crawl measured 43 KB and finished in under a minute, and pricing.billingMode is metered_result, so cost scales with what it finds — set limit. Send a fresh Idempotency-Key per distinct crawl; reusing one returns the earlier job instead of starting a new one. For a handful of known URLs post_firecrawl_batch_scrape is cheaper, and for structure alone post_firecrawl_map costs far less.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesThe HTTPS root URL to crawl. PDF URLs are not supported.
delayNoDelay in seconds between requests (0 to 30).
limitYesMaximum number of pages to crawl.
sitemapNoHow the site's sitemap is used during discovery.
excludePathsNoSkip URLs whose path matches one of these patterns.
includePathsNoOnly crawl URLs whose path matches one of these patterns.
scrapeOptionsNoPer-page scrape options applied while crawling. On the metered profile output is always markdown.
maxConcurrencyNoMaximum number of concurrent page fetches (1 to 20).
Idempotency-KeyYesUnique key (1 to 191 characters) that makes the submit idempotent. Re-submitting with the same key returns the original job.
allowSubdomainsNoFollow links into subdomains of the root domain.
crawlEntireDomainNoCrawl the whole domain rather than only the subtree under the root URL.
maxDiscoveryDepthNoMaximum link-discovery depth from the root URL.
allowExternalLinksNoFollow links to external domains.
ignoreQueryParametersNoTreat URLs that differ only by query string as the same page.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observed

TDQS

A4.9/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

The description adds substantial behavior beyond the annotations: it reveals the async lifecycle (HTTP 202, job envelope with `id`/`status`/`pricing`, `output` null initially), the cost model (`pricing.billingMode` is `metered_result`, cost scales with what the crawl finds), and idempotency semantics (reusing an `Idempotency-Key` returns the earlier job). None of this is derivable from `readOnlyHint`/`openWorldHint`/`idempotentHint` alone, and it does not contradict any annotation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

While long, the description is densely functional and every sentence earns its place: purpose first, then requirements, async flow, cost caveat, idempotency rule, and alternatives. The most decision-critical facts (required params, async nature, cost risk, when to use a sibling) are front-loaded, and there is no filler or restatement of what the input schema already says.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex tool (14 parameters, nested objects, async job lifecycle), the description is complete: it covers the job envelope, output shape, polling statuses, the cost implications of the metered billing mode, the idempotency key requirement, and the correct alternative tools. The input schema covers parameter semantics and constraints, and the output schema exists for return values, so no essential guidance is left to inference.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, which sets the baseline at 3. The description adds meaning beyond the schema by flagging `limit` as a cost control ("cost scales with what it finds — set `limit`") and by grouping `includePaths`, `excludePaths`, `maxDiscoveryDepth`, `crawlEntireDomain`, and `allowSubdomains` as steering knobs, plus `delay` and `maxConcurrency`. It doesn't elaborate every parameter, but the schema already documents each one, so the added value pushes it a notch above baseline.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The first sentence states a specific verb and resource — "Crawl a whole site rooted at `url`" — plus the expected outcome ("content of every page it keeps"). It also differentiates from siblings by naming `post_firecrawl_batch_scrape` (better for a handful of known URLs) and `post_firecrawl_map` (better for structure alone), so an agent can pick the right tool without opening other schemas.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells the agent when to use this tool versus alternatives: "For a handful of known URLs `post_firecrawl_batch_scrape` is cheaper, and for structure alone `post_firecrawl_map` costs far less." It also gives a concrete usage recipe (required `url`, `limit`, and `Idempotency-Key`; steering knobs like `includePaths` and `excludePaths`) and states the full polling lifecycle with the sibling to poll (`get_firecrawl_crawl_job`) and the terminal statuses to wait for.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources