MCP Server Firecrawl
Firecrawl MCP 服务器
使用 Firecrawl API 进行网页抓取、内容搜索、站点爬取和数据提取的模型上下文协议 (MCP) 服务器。
特征
网页抓取:使用可自定义的选项从任何网页提取内容
移动设备模拟
广告和弹出窗口拦截
内容过滤
结构化数据提取
多种输出格式
内容搜索:智能搜索功能
多语言支持
基于位置的结果
可定制的结果限制
结构化输出格式
网站爬取:高级网页爬取功能
深度控制
路径过滤
速率限制
进度追踪
网站地图集成
站点映射:生成站点结构图
子域名支持
搜索过滤
链接分析
视觉层次
数据提取:从多个 URL 中提取结构化数据
架构验证
批处理
网络搜索丰富
自定义提取提示
Related MCP server: MCP Firecrawl Server
安装
# Global installation
npm install -g @modelcontextprotocol/mcp-server-firecrawl
# Local project installation
npm install @modelcontextprotocol/mcp-server-firecrawl快速入门
从开发者门户获取您的 Firecrawl API 密钥
设置您的 API 密钥:
Unix/Linux/macOS(bash/zsh):
export FIRECRAWL_API_KEY=your-api-keyWindows(命令提示符):
set FIRECRAWL_API_KEY=your-api-keyWindows(PowerShell):
$env:FIRECRAWL_API_KEY = "your-api-key"替代方案:使用 .env 文件(推荐用于开发):
# Install dotenv npm install dotenv # Create .env file echo "FIRECRAWL_API_KEY=your-api-key" > .env然后在你的代码中:
import dotenv from 'dotenv'; dotenv.config();运行服务器:
mcp-server-firecrawl
一体化
克劳德桌面应用程序
添加到您的 MCP 设置:
{
"firecrawl": {
"command": "mcp-server-firecrawl",
"env": {
"FIRECRAWL_API_KEY": "your-api-key"
}
}
}Claude VSCode 扩展
添加到您的 MCP 配置:
{
"mcpServers": {
"firecrawl": {
"command": "mcp-server-firecrawl",
"env": {
"FIRECRAWL_API_KEY": "your-api-key"
}
}
}
}使用示例
网页抓取
// Basic scraping
{
name: "scrape_url",
arguments: {
url: "https://example.com",
formats: ["markdown"],
onlyMainContent: true
}
}
// Advanced extraction
{
name: "scrape_url",
arguments: {
url: "https://example.com/blog",
jsonOptions: {
prompt: "Extract article content",
schema: {
title: "string",
content: "string"
}
},
mobile: true,
blockAds: true
}
}网站抓取
// Basic crawling
{
name: "crawl",
arguments: {
url: "https://example.com",
maxDepth: 2,
limit: 100
}
}
// Advanced crawling
{
name: "crawl",
arguments: {
url: "https://example.com",
maxDepth: 3,
includePaths: ["/blog", "/products"],
excludePaths: ["/admin"],
ignoreQueryParameters: true
}
}站点地图
// Generate site map
{
name: "map",
arguments: {
url: "https://example.com",
includeSubdomains: true,
limit: 1000
}
}数据提取
// Extract structured data
{
name: "extract",
arguments: {
urls: ["https://example.com/product1", "https://example.com/product2"],
prompt: "Extract product details",
schema: {
name: "string",
price: "number",
description: "string"
}
}
}配置
有关详细的设置选项,请参阅配置指南。
API 文档
有关详细的端点规范,请参阅API 文档。
发展
# Install dependencies
npm install
# Build
npm run build
# Run tests
npm test
# Start in development mode
npm run dev示例
查看示例目录以获取更多使用示例:
基本抓取: scrape.ts
爬取和映射: crawl-and-map.ts
错误处理
服务器实现了强大的错误处理:
使用指数退避算法进行速率限制
自动重试
详细错误消息
调试日志记录
安全
API 密钥保护
请求验证
域名允许列表
速率限制
安全错误消息
贡献
请参阅CONTRIBUTING.md了解贡献指南。
执照
MIT 许可证 - 详情请参阅许可证。
Available Tools
2 toolsextractC
Extracts structured data from URLs
| Name | Required | Description | Default |
|---|---|---|---|
| urls | Yes | URLs to extract from | |
| prompt | No | Extraction guidance prompt | |
| schema | No | Data structure schema | |
| ignoreSitemap | No | Ignore sitemap.xml during processing | |
| enableWebSearch | No | Use web search for additional data | |
| includeSubdomains | No | Include subdomains in processing |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations exist, so the description must disclose behavior. It does not mention network requests, rate limits, authentication, error handling, or whether it is idempotent. Only states 'extracts', which is vague.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Very short at 7 words, but under-specified. Lacks structure such as sections or examples. Conciseness would be appropriate if complete, but here it is insufficient.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the complexity (6 parameters, nested objects, sibling tools, no output schema), the description is extremely incomplete. No information on return format, usage patterns, or edge cases.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description does not add any extra meaning beyond the schema's parameter descriptions. No elaboration on how parameters interact.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states the tool extracts structured data from URLs, but does not differentiate from sibling 'map'. The verb and resource are clear, but the scope is vague as 'structured data' is not defined.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance on when to use this tool versus alternatives like 'map', no prerequisites or exclusions provided. The description is insufficient for an agent to decide usage context.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
mapC
Maps a website's structure
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Base URL to map | |
| limit | No | Maximum links to return | |
| search | No | Search query for mapping | |
| timeout | No | Request timeout | |
| sitemapOnly | No | Only use sitemap.xml for mapping | |
| ignoreSitemap | No | Ignore sitemap.xml during mapping | |
| includeSubdomains | No | Include subdomains in mapping |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are present, so the description must disclose behavioral traits. It only says 'Maps a website's structure' with no details on whether it crawls, destructiveness, rate limits, or return format. This is insufficient.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, short sentence, which is efficient but under-specified for a tool with 7 parameters. It could include key details without sacrificing conciseness.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's complexity (7 parameters, no output schema, no annotations), the description lacks essential context such as return values, process details, or limitations. It is highly incomplete.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema has 100% description coverage for all 7 parameters, so the baseline is 3. The description adds no additional meaning beyond the schema, making it adequate but not improved.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description states a specific verb ('Maps') and resource ('a website's structure'), making the core purpose clear. However, it does not differentiate from the sibling tool 'extract', which could lead to ambiguity.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives (e.g., 'extract'). The description lacks context for selection or prerequisites.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v1.0.1- First observed
extract - First observed
map
TDQS
Scored across 2 tools
The two tools have clearly distinct purposes: one extracts structured data from URLs, the other maps a website's structure. There is no overlap in functionality.
Both tool names are single imperative verbs ('extract' and 'map'), following a consistent and concise naming pattern.
With only 2 tools, the server is on the lower end of acceptable scope. It covers basic functionality but feels minimal for a web crawling service.
The server provides core data extraction and site mapping, but lacks advanced features like recursive crawling or search, which are notable gaps for the domain.
Maintenance
Related MCP Connectors
MCP server for Firecrawl — web search, scraping, and biomedical/arXiv paper search.
Firecrawl MCP — wraps the Firecrawl API (firecrawl.dev) for web
Zenrows MCP server — Fetch, Extract, Batch, and Browser Sessions for AI coding assistants
Related MCP Servers
- FlicenseCqualityCmaintenanceBuilt as a Model Context Protocol (MCP) server that provides advanced web search, content extraction, web crawling, and scraping capabilities using the Firecrawl API.41-
- FlicenseDqualityDmaintenanceA server that provides tools to scrape websites and extract structured data from them using Firecrawl's APIs, supporting both basic website scraping in multiple formats and custom schema-based data extraction.23-
- AlicenseBqualityDmaintenanceA Model Context Protocol server that enables AI assistants to perform advanced web scraping, crawling, searching, and data extraction through the Firecrawl API.9117,089 npmMIT
- AlicenseAqualityDmaintenanceA Model Context Protocol server that enables web scraping, crawling, and content extraction capabilities through integration with Firecrawl.8117,089 npm2MIT