Skip to main content
Glama

firecrawl_extract

Extract structured data from one or more URLs with an LLM, then poll status to retrieve results. Use it to turn web pages into schema-defined records for marketing workflows.

Instructions

Extract structured data from one or more URLs using an LLM. Poll results with extract_status.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlsYesThe URLs to extract data from. URLs should be in glob format.
promptNoPrompt to guide the extraction process.
schemaNoSchema to define the structure of the extracted data. Must conform to JSON Schema.
showSourcesNoWhen true, the sources used to extract the data will be included in the response as `sources`.
ignoreSitemapNoWhen true, sitemap.xml files will be ignored during website scanning.
scrapeOptionsNoOptions applied when scraping pages for extraction.
enableWebSearchNoWhen true, the extraction will use web search to find additional data.
threatProtectionNoPer-request threat protection override. Enterprise feature.
ignoreInvalidURLsNoIf invalid URLs are specified, they are ignored and returned in invalidURLs instead of failing the request. The server applies true when this is absent.
includeSubdomainsNoWhen true, subdomains of the provided URLs will also be scanned. The server applies true when this is absent.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.5/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations (readOnlyHint=false, openWorldHint=true) are minimal and do not reveal the async job lifecycle, so the description correctly surfaces the most important behavioral trait: results must be polled via extract_status. It stops there, omitting what the initial call returns (a job id?), failure behavior, and whether web-search/credit costing applies. Decent added context over annotations, but incomplete.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tight sentences with zero waste. The core purpose is front-loaded and the polling instruction follows immediately. Nothing to trim and nothing misplaced.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With 10 parameters, no output schema, and an async job model, the description covers the poll step but not the full lifecycle: it does not explain what the initial invocation returns or how to correlate extract_status results back to it. For a complex async tool with no output schema to fall back on, this is a real gap.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all 10 parameters (urls, prompt, schema, scrapeOptions, enableWebSearch, etc.) are already documented in the schema. The description adds no parameter-level detail (e.g., the glob URL format or JSON Schema requirement) beyond what the schema already provides. Baseline 3 is appropriate.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb+resource ('Extract structured data from one or more URLs') and adds the key mechanism ('using an LLM'), so the agent understands this is LLM-driven structured extraction rather than raw scraping. It does not, however, distinguish itself from the very close sibling firecrawl_scrape, leaving the boundary between the two to inference.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The second sentence ('Poll results with extract_status') is genuine usage guidance: it tells the agent this is an async operation that must be polled. But it gives no guidance on when to choose this tool over firecrawl_scrape or other extraction options, and does not name any exclusion criteria. Usage is implied rather than fully specified.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.