Firecrawl MCP Server
Server Details
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
- Status
- Healthy
- Last Tested
- Transport
- Streamable HTTP · MCP 2025-11-25
- URL
- Repository
- firecrawl/firecrawl-mcp-server
- GitHub Stars
- 7,522
- Server Listing
- Firecrawl MCP Server
TDQS
Scored across 3 tools
The three tools target distinct primary actions: parsing supported documents, scraping web URLs, and searching the web. Parse vs. scrape could be confused, but the descriptions explicitly separate local/hosted documents from remote URLs.
All tools use the same firecrawl_ prefix followed by a clear lowercase verb (parse, scrape, search). The naming is fully consistent and predictable.
Three tools is thin for a Firecrawl server whose descriptions reference map, crawl, find_tools, and research capabilities. The core parse/scrape/search surface is useful but under-scoped for the apparent broader purpose.
The descriptions repeatedly mention firecrawl_map, firecrawl_crawl, firecrawl_find_tools, and firecrawl_research_* tools that are not included, creating dead references. Agents cannot perform common lifecycle operations like site mapping, crawling, or research discovery with the available surface.
Available Tools
3 toolsfirecrawl_parseARead-onlyInspect
Parse one supported document into markdown, HTML, links, summary, targeted answers, or JSON matching a schema. Supported inputs include common HTML, PDF, Word, RTF, OpenDocument, and spreadsheet files; PDF parsing can be bounded with pdfOptions.maxPages.
Local MCP reads filePath from the server filesystem. Hosted MCP uses two calls: first provide filePath to receive upload instructions, upload locally, then call again with the returned uploadRef; do not send both fields together. Remote web URLs belong in firecrawl_scrape.
Set redactPII to request redaction of personally identifiable information in the returned content. zeroDataRetention requires an eligible authenticated account; omit it for anonymous keyless use. Returns upload instructions for hosted phase one or parsed document content for the final call. Authenticated final responses can include a data.metadata.scrapeId for optional parse feedback.
| Name | Required | Description | Default |
|---|---|---|---|
| proxy | No | ||
| maxAge | No | Ignored: parse never reuses or stores indexed content. | |
| formats | No | ||
| parsers | No | ||
| filePath | No | Phase 1 only: path to the local file on the caller/harness machine. Hosted MCP will not read or stat this path; it is used only to produce upload instructions. | |
| redactPII | No | ||
| uploadRef | No | Phase 2 only: short-lived upload reference returned by phase 1 after the local PUT upload completes. | |
| pdfOptions | No | ||
| contentType | No | Phase 1 MIME type override. If omitted, the server infers it from the file extension without reading the file. | |
| excludeTags | No | ||
| includeTags | No | ||
| jsonOptions | No | ||
| queryOptions | No | ||
| storeInCache | No | ||
| onlyMainContent | No | ||
| declaredSizeBytes | No | Optional phase 1 size declaration. Hosted MCP does not stat the file; provide this only if the caller already knows it. | |
| zeroDataRetention | No | ||
| removeBase64Images | No | ||
| skipTlsVerification | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| raw | No | Response body when it was not JSON. |
| data | No | Parsed document content; can include `data.metadata.scrapeId` for parse feedback. |
| mode | No | Which phase of the hosted flow produced this response. |
| error | No | Error message or error object when the call did not succeed. |
| notes | No | Hosted phase one: constraints on completing the upload flow. |
| upload | No | Hosted phase one: how to upload the local file. |
| message | No | Guidance for the next call. |
| success | No | Whether the API call succeeded. |
| warning | No | Non-fatal warning about the result. |
| agent_hints | No | Optional response guidance from the Firecrawl API. |
| nextToolCall | No | Hosted phase one: the second `firecrawl_parse` call to make once the upload succeeds, as `{name, arguments}`. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations already establish a safe read-only, non-destructive, closed-world operation, and the description adds substantial context on top: the two-phase upload behavior, what each call returns, PII redaction semantics, the zeroDataRetention account requirement, and the optional scrapeId feedback hook. None of this is derivable from the annotations.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Front-loaded with purpose and formats, then mechanics, then caveats, with no filler sentences. The one return-value sentence is arguably redundant given the output schema, but it is short and useful for disambiguating the two-phase flow.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 19-parameter tool with nested objects, the description nails the invocation flow and a few key flags but omits guidance on the options that govern the requested output modes (json/query schemas, tag filtering, formats/parsers). The presence of an output schema excuses return-value detail, but the parameter gaps keep this at minimum-viable.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
With only 26% schema description coverage across 19 parameters, the description must compensate more than it does. It meaningfully explains redactPII, zeroDataRetention, pdfOptions.maxPages, filePath and uploadRef, but leaves format-driving options (jsonOptions, queryOptions, includeTags/excludeTags, formats, parsers) undocumentable in either place.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb and resource ('Parse one supported document') and enumerates the output formats and the supported input file types. It also distinguishes itself from the sibling firecrawl_scrape by explicitly routing remote web URLs there.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Gives explicit operating conditions: local MCP reads filePath from the server filesystem, hosted MCP requires a two-call upload handshake with the warning not to send both fields together, and remote URLs belong in firecrawl_scrape. It also states the auth precondition for zeroDataRetention and tells anonymous users to omit it.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firecrawl_scrapeAInspect
Scrape one URL and return its content: markdown by default, or HTML, links, screenshots, branding data, a targeted answer, or JSON matching a supplied schema. Use it when the request identifies a page and needs its content or defined fields. Use firecrawl_search when additional web sources are needed; on an authenticated session, firecrawl_map lists a site's URLs and firecrawl_crawl collects a set of pages.
Firecrawl may serve recently indexed content; set maxAge: 0 for a live fetch or a smaller maxAge to bound staleness. A successful response does not by itself confirm the page is still current. Browser actions can change the live page when interactive actions are enabled. Authenticated responses can include a metadata.scrapeId for optional scrape feedback.
On an authenticated session with Alexandria access, firecrawl_search with sources unset and firecrawl_find_tools can discover providers for the same fields across several pages; a matching provider returns typed records in one call. Keyless sessions have no provider matches.
Alexandria mode, on an authenticated session with Alexandria access: alexandria selects catalogued capability execution and is mutually exclusive with url.
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | ||
| proxy | No | ||
| maxAge | No | ||
| mobile | No | ||
| formats | No | ||
| parsers | No | ||
| profile | No | ||
| timeout | No | Execution timeout in milliseconds. | |
| waitFor | No | ||
| location | No | ||
| lockdown | No | ||
| redactPII | No | ||
| requestId | No | Idempotency key bound to one Alexandria execution payload. Generated when omitted and returned with the result. | |
| alexandria | No | Catalogued Alexandria capability invocation, mutually exclusive with url. One {provider, capability, options} object or an array of 1-10, with contracts available through firecrawl_search or firecrawl_find_tools. Each call may include version to pin a published workflow; omitting it uses latest. Only requestId and timeout are supported alongside alexandria. The selected contract marks required inputs and any requiresOneOf groups (at least one member per group); it may include example.request/example.response and response.key (which may differ from records). Where pagination is declared, its fields govern paging with the same filters; catalogue next is separate from provider pagination. Returns per-capability results in data.alexandria with data, records, or an error with a code; individual capabilities can fail even when the outer response succeeds. Requires an API key on a team with Alexandria enabled. Some providers require accepted terms; blocked requests return the applicable requirements. | |
| pdfOptions | No | ||
| toolDetail | No | URL mode only: domain discovery detail, summary by default; compact returns provider/capability/description, full includes contracts. Ignored with alexandria. | |
| domainTools | No | URL mode only: include domain-matched Alexandria tools for the page in tools on the returned document. Ignored with alexandria. | |
| excludeTags | No | ||
| includeTags | No | ||
| jsonOptions | No | ||
| queryOptions | No | ||
| storeInCache | No | ||
| onlyMainContent | No | ||
| screenshotOptions | No | ||
| zeroDataRetention | No | ||
| removeBase64Images | No | ||
| skipTlsVerification | No |
Output Schema
| Name | Required | Description |
|---|---|---|
| data | No | Alexandria mode: per-capability results in `data.alexandria`, each with `data`, `records`, or an `error`. |
| html | No | Processed HTML of the page. |
| json | No | Structured data matching the requested JSON schema or prompt. |
| menu | No | Menu data extracted from the page. |
| audio | No | Audio extracted from the page. |
| error | No | Error message or error object when the call did not succeed. |
| links | No | Links found on the page. |
| pages | No | Physical PDF pages, when `parsers[].pages` is set. |
| tools | No | Domain-matched Alexandria tools for the page, when `domainTools` is set. |
| video | No | Video extracted from the page. |
| answer | No | Targeted answer to the question that was asked of the page. |
| blocks | No | Typed PDF layout blocks, when `parsers[].blocks` is set. |
| images | No | Images found on the page. |
| actions | No | Results of the browser actions that ran during the scrape. |
| message | No | Guidance that accompanies the result. |
| product | No | Product data extracted from the page. |
| rawHtml | No | Unprocessed HTML of the page. |
| receipt | No | Billing receipt for the execution. |
| success | No | Whether the API call succeeded. |
| summary | No | Summary of the page content. |
| warning | No | Non-fatal warning about the result. |
| branding | No | Branding data extracted from the page. |
| delivery | No | `retained` when the full result stayed server-side instead of being inlined. |
| markdown | No | Page content as markdown. |
| metadata | No | Page metadata; authenticated responses can include `metadata.scrapeId` for scrape feedback. |
| nextTool | No | A follow-up tool call (`{name, arguments}`) that continues or inspects this result. |
| requestId | No | Identifier of this Alexandria execution. |
| scrape_id | No | Identifier of the underlying scrape. |
| attributes | No | Values collected by the requested attribute selectors. |
| highlights | No | Highlighted passages from the page. |
| screenshot | No | Screenshot of the page. |
| agent_hints | No | Optional response guidance from the Firecrawl API. |
| creditsCost | No | Credits this call consumed. |
| workspaceId | No | Workspace holding a retained result, for inspection through virtual Bash. |
| feedbackTool | No | Pointer to the feedback tool for reporting how this result served the task. |
| responseBytes | No | Size of the full result in bytes. |
| changeTracking | No | Change-tracking comparison against the previous scrape. |
| idleTtlSeconds | No | Seconds a retained workspace stays available while idle. |
| estimatedTokens | No | Estimated token cost of the full result. |
| inlineTokenBudget | No | Token budget above which a result is retained rather than inlined. |
| tokenEstimateMethod | No | How `estimatedTokens` was derived. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations declare readOnlyHint=false, destructiveHint=false, openWorldHint=true, but the description adds genuinely non-obvious behavior: cached/recently-indexed content and the maxAge=0 escape hatch, the caveat that a successful response does not confirm freshness, that browser actions can mutate the live page, scrapeId feedback, and keyless-session limitations. It stops short of describing pagination or partial-failure behavior for URL mode, so it is not a 5.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The first paragraph is well front-loaded and earns its place, but the second and third paragraphs are dense run-on sentences with heavy qualifiers ('On an authenticated session with Alexandria access, firecrawl_search with sources unset and firecrawl_find_tools can discover providers...'). The content is relevant but the structure makes it hard to scan.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given a 27-parameter tool with nested objects and an output schema, the description covers the important mode distinctions, freshness semantics, auth requirements, and the alexandria-exclusive mode. Return values are covered by the output schema. The main gap is the undocumented parameter surface, which is partly excusable given the volume.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is only 19% across 27 parameters, so the description carries a heavy burden. It glosses the format options and explains maxAge and the alexandria/url mutual exclusion, but leaves the majority of parameters (proxy, location, parsers, profile, screenshotOptions, jsonOptions, redactPII, lockdown, and more) undocumented in both places.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
Opens with a specific verb+resource ('Scrape one URL') and enumerates the concrete output formats the caller can request. It explicitly separates itself from firecrawl_search, firecrawl_map, and firecrawl_crawl by naming the condition that selects each.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
States the selection condition ('when the request identifies a page and needs its content or defined fields') and names the alternatives with their own triggers. The Alexandria paragraph adds a clear mutually-exclusive usage path with its own precondition (authenticated session with Alexandria access).
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
firecrawl_searchARead-onlyInspect
Search web, news, or image sources and return ranked results with query-relevant highlights. Each web result is a title, URL, and description; use firecrawl_scrape on a result URL when the excerpt is not enough.
Authenticated search also returns matching Alexandria data providers in data.tools (companies, people, jobs, finance and filings, public records and government spending, real estate, places and restaurants, retail and prices, package registries and developer data, news, research, and more). Prefer a provider over scraping pages when the task needs the same fields across several entities, exact figures or timestamps, provenance, or many records; use web results when they already answer the question. A search with sources: ["web"] omits semantic provider discovery; domainTools: true can still return website-matched tools. Web-only results use domainTools: false.
On an authenticated session, tool matches describe available capabilities; firecrawl_find_tools returns their contracts and firecrawl_scrape with an alexandria body executes a selected capability. Keyless sessions get no Alexandria matches in data.tools.
For a programming question, add categories: ["developer"]; its hits return in data.web with category: "developer". categories: ["research"] restricts web results to research-affiliated websites; the firecrawl_research_* tools are a separate surface over paper abstracts and full text (PubMed, bioRxiv, medRxiv, arXiv). Query operators, domain filters, categories, toolDetail and scrapeOptions are described on their parameters. Returns source-type result groups and usage metadata. Authenticated responses can include an id for optional search feedback.
| Name | Required | Description | Default |
|---|---|---|---|
| tbs | No | ||
| limit | No | ||
| query | Yes | Query for web and semantic tool discovery. Operators include quoted phrases, `-term`, `site:host`, `inurl:term`, `intitle:term`, and `related:host`; the set is non-exhaustive. Catalogue browsing is available through firecrawl_find_tools. | |
| filter | No | ||
| sources | No | Search sources; authenticated sessions default to web + alexandria, keyless sessions to web only. A search with sources: ["web"] omits semantic provider discovery; domainTools: true can still return website-matched tools. Web-only results use domainTools: false. Use ["alexandria"] alone for provider discovery without web results. | |
| location | No | ||
| categories | No | Limit results to specific source types. `research` restricts ordinary web results to research-affiliated websites and returns page snippets, which is separate from the `firecrawl_research_*` tools that search paper abstracts and full text across biomedical (PubMed, bioRxiv, medRxiv) and arXiv literature; `pdf` searches PDF results; `developer` searches an index built for coding agents over public repositories, GitHub issues, merged pull requests, repository READMEs, and code documentation. `developer` returns hits in `data.web` with `category: "developer"`; the other categories also filter `data.web`. | |
| enterprise | No | ||
| highlights | No | Return query-relevant page excerpts for web and news results when available (default). Highlights appear in web `description` and news `snippet`; otherwise, original snippets are returned. Set to false to keep the original search snippets. | |
| toolDetail | No | Compact by default. Compact returns only provider, capability and description; full includes contracts. Inspect selected compact tools with firecrawl_find_tools providers and capabilities. | |
| domainTools | No | Include domain-matched tools for result URLs. Defaults to true when Alexandria is combined with web, news or images; semantic-only search leaves domain matching off. | |
| scrapeOptions | No | Attach page content for web results in the same call. These fetches ignore maxAge, so use firecrawl_scrape when you need a live fetch. scrapeOptions fetches web pages, never Alexandria provider tools. | |
| excludeDomains | No | Hostnames to leave out of results. Mutually exclusive with includeDomains. | |
| includeDomains | No | Hostnames to restrict results to. Mutually exclusive with excludeDomains. |
Output Schema
| Name | Required | Description |
|---|---|---|
| id | No | Search identifier, for optional `firecrawl_search_feedback`. |
| data | No | Ranked results grouped by source, such as `web`, `news`, `images`, and `alexandria`. |
| error | No | Error message or error object when the call did not succeed. |
| tools | No | Domain-matched Alexandria tools for the results. |
| success | No | Whether the API call succeeded. |
| warning | No | Non-fatal warning about the result. |
| nextTool | No | A follow-up tool call that continues this search. |
| agent_hints | No | Optional response guidance from the Firecrawl API. |
| creditsUsed | No | Credits this search consumed. |
| feedbackTool | No | Pointer to the feedback tool for this search. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations cover read-only/open-world safety, but the description goes further: keyless vs authenticated behavior diverges (no Alexandria matches), default sources differ by session, scrapeOptions fetches ignore maxAge, and an optional search-feedback id is returned. These auth- and session-dependent traits are not derivable from annotations or schema.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Dense but mostly front-loaded and every paragraph carries routing or auth information. Some Alexandria/tool-discovery detail is repeated across sentences (toolDetail vs firecrawl_find_tools appears twice), costing a little tightness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
For a 14-parameter, nested-schema, open-world search tool with an output schema, the description covers session-dependent defaults, provider-vs-scrape tradeoffs, category semantics, and return grouping without redundantly explaining the document-return schema.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is only 64%, and the description meaningfully compensates for several under-documented params (sources defaults, domainTools defaults and web-only behavior, toolDetail compact/full semantics, categories including developer/research). It does defer operators and scrapeOptions to the parameter docs, which is reasonable since those are well-described in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb (search) and resource (web, news, image sources) plus the return shape (ranked results with highlights). It also distinguishes itself from firecrawl_scrape by naming when to use that sibling instead of this one.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Explicit routing rules: use firecrawl_scrape when the excerpt is insufficient, prefer an Alexandria provider when the task spans several entities or needs exact figures, use web results when they already answer. Also explains sources/domainTools interactions and category selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
3 tool updates
- First observed
firecrawl_parse - First observed
firecrawl_scrape - First observed
firecrawl_search
Related MCP Connectors
Scrape, crawl and search the web for AI agents via MCP.
Docs: https://docs.keenable.ai/mcp-server Keenable is a free, remote MCP server that gives agents access to the web index. Search the web with ranked results and date/site filters, then fetch any indexed page as clean markdown. Works out of the box with no account or API key.
Firecrawl MCP — wraps the Firecrawl API (firecrawl.dev) for web
Related MCP Servers
- AlicenseNot gradedqualityBmaintenanceWeb scraping and search MCP server that wraps Firecrawl API for URL discovery and web search with optional content retrieval.140 npm1MIT
- AlicenseAqualityCmaintenanceMCP server for web page fetching (converting to Markdown/text with automatic fallback between Tavily and Firecrawl) and web search via Tavily.2MIT
- AlicenseNot gradedqualityDmaintenanceA Model Context Protocol (MCP) server implementation that integrates with Firecrawl for web scraping capabilities.190,664 npm1MIT
- FlicenseCqualityCmaintenanceBuilt as a Model Context Protocol (MCP) server that provides advanced web search, content extraction, web crawling, and scraping capabilities using the Firecrawl API.41-
Glama MCP Gateway
Add one secure layer between your agents and this server.