Skip to main content
Glama
stevenzxs

Crawl4AI MCP Server

by stevenzxs
README.md
# Crawl4AI MCP Server

本地运行的 Crawl4AI MCP Server,为 AI 助手提供强大的网页爬取能力。

## ✨ 功能特性

- **🌐 网页爬取**:爬取单个或多个 URL,返回干净的 Markdown
- **📝 结构化提取**:使用 CSS/XPath 或 LLM 提取结构化数据  
- **📸 截图功能**:获取网页截图
- **🕸️ 深度爬取**:支持深度爬取整个网站
- **🚀 本地运行**:完全本地运行,无需 API Key(基础功能)
- **⚡ 高性能**:异步并发,智能缓存

## 🚀 快速开始

### 1. 安装

```bash
# 克隆项目
git clone <your-repo-url>
cd crawl4ai-mcp

# 使用启动脚本(推荐)
chmod +x start.sh
./start.sh

# 或者手动安装
uv sync
uv run crawl4ai-setup
```

### 2. 配置

```bash
# 复制环境变量示例
cp .env.example .env

# 编辑 .env 文件,添加你的 API keys(用于结构化提取)
```

### 3. 运行

```bash
# stdio 模式(用于 Claude Code)
uv run crawl4ai-mcp

# HTTP 模式(用于开发调试)
uv run crawl4ai-mcp --transport http --port 8000
```

## 🔧 配置 Claude Code

### stdio 模式(推荐)

```bash
claude mcp add crawl4ai uv run --project /path/to/crawl4ai-mcp crawl4ai-mcp
```

### HTTP 模式

```bash
# 1. 启动服务器
uv run crawl4ai-mcp --transport http --port 8000

# 2. 添加到 Claude Code
claude mcp add --transport http crawl4ai http://localhost:8000/mcp
```

## 📚 可用工具

### crawl_url - 爬取网页

```python
crawl_url(
    url="https://example.com",
    word_count_threshold=10,
    bypass_cache=False,
    magic=False
)
```

### crawl_multiple - 批量爬取

```python
crawl_multiple(
    urls=["https://example.com/page1", "https://example.com/page2"],
    max_concurrent=3,
    word_count_threshold=10
)
```

### extract_structured - 结构化提取

```python
extract_structured(
    url="https://example.com/products",
    instruction="提取所有产品名称和价格",
    provider="openai/gpt-4o-mini",
    api_token="your-api-key"
)
```

### get_screenshot - 网页截图

```python
get_screenshot(
    url="https://example.com",
    full_page=True,
    viewport_width=1920,
    viewport_height=1080
)
```

### deep_crawl - 深度爬取

```python
deep_crawl(
    url="https://example.com",
    max_depth=2,
    max_pages=10,
    strategy="bfs"  # 或 "dfs"
)
```

## 📖 文档

- [详细使用指南](GUIDE.md) - 完整的工具说明和使用示例
- [Crawl4AI 文档](https://docs.crawl4ai.com/) - 底层库文档
- [MCP 协议](https://modelcontextprotocol.io/) - MCP 协议文档

## 🛠️ 开发

```bash
# 运行测试
make test

# 代码格式化
make fmt

# 代码检查
make lint

# 类型检查
make typecheck

# 运行所有检查
make check
```

## 🔒 环境变量

```bash
# LLM API Keys(用于结构化提取)
OPENAI_API_KEY=sk-...
ANTHROPIC_API_KEY=sk-ant-...

# 或使用通用配置
LLM_PROVIDER=openai/gpt-4o-mini
LLM_API_KEY=sk-...
```

## 📄 许可证

MIT License

## 🤝 贡献

欢迎提交 Issue 和 PR!

## 💬 支持

- [Crawl4AI Discord](https://discord.gg/jP8KfhDhyN)
- [MCP 社区](https://github.com/modelcontextprotocol)

Maintenance

ActivityInactive
ResponsivenessNo issues