Crawl4AI MCP
You can use this MCP server to let AI assistants crawl websites, extract Markdown content, and manage crawl results.
Crawl sites with configurable
max_depth,max_pages, and optional external links.Fetch a single page directly without following links or saving a file (
crawl_page).Save results as Markdown and expose them as MCP resources under
crawl://results/....Return extracted content in responses, with full content saved to file if truncated.
Use CSS selectors,
wait_for_selector, and delays for targeted or dynamic/SPA extraction.Optionally enable magic mode, custom JavaScript (
CRAWL4AI_MCP_ALLOW_JS=true), and persistent sessions (session_id/close_session).Get progress notifications and detailed crawl statistics, including skipped pages.
Keep crawling secure by default: public HTTP(S) only, no private networks, no JS unless opted in.
Configure results directory, timeouts, concurrency, session TTL, logging, and verbose output.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@Crawl4AI MCPcrawl https://example.com and summarize the main content"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Web Crawler MCP
A powerful web crawling tool that integrates with AI assistants via the MCP (Model Context Protocol). This project allows AI assistants to crawl websites, extract dynamic content, navigate through links, and save structured Markdown files directly.
📋 Features
Native integration with AI assistants via MCP
Return scraped Markdown content directly to the AI
Extracts and surfaces internal/external links for AI navigation
Website crawling with configurable depth and page limit
Detailed crawl result statistics, including the list of skipped pages
Live progress notifications during long crawls
Saved results exposed as MCP resources (
crawl://results/...)Error and not found page handling
Secure by default: only public http(s) URLs, JavaScript disabled unless you opt in
Advanced Scraping Capabilities:
Magic Mode: crawl4ai heuristics that simulate a real browser and help with some anti-bot protections (not a guaranteed bypass)
Targeted Extraction: Fetch only what you need using CSS selectors
Custom JavaScript (opt-in): Execute code before extraction (clicks, scrolls, form fills)
Persistent Sessions: One browser is shared across calls, so a
session_idkeeps cookies and state until it is closed (close_session) or stays idle too longSPA Support: Wait for dynamic CSS selectors or set explicit pre-extraction delays
Related MCP server: Scrapling MCP Server
🚀 MCP Configuration
The simplest and recommended way to use this tool is via uvx. The @latest suffix makes uvx check PyPI at each start and run the newest published version (without it, uvx keeps reusing the version it cached the first time).
Prerequisites
uv installed on your system.
Setup for AI Assistants (e.g., Claude Desktop, Cline)
Add the following to your AI Assistant's MCP configuration file (e.g., cline_mcp_settings.json or claude_desktop_config.json):
Note: Python 3.12 or 3.13 is required (crawl4ai does not support 3.14 yet). Specifying
--python 3.13is recommended, especially on Windows, to avoid compilation issues with certain dependencies.
From PyPI (recommended):
{
"mcpServers": {
"crawl": {
"command": "uvx",
"args": [
"--python",
"3.13",
"crawl4ai-mcp-llm@latest"
],
"disabled": false,
"autoApprove": [],
"timeout": 600
}
}
}From GitHub (latest unreleased):
{
"mcpServers": {
"crawl": {
"command": "uvx",
"args": [
"--python",
"3.13",
"--from",
"git+https://github.com/laurentvv/crawl4ai-mcp-llm",
"crawl4ai-mcp-llm"
],
"disabled": false,
"autoApprove": [],
"timeout": 600
}
}
}Claude Code:
claude mcp add crawl -s user -- uvx --python 3.13 crawl4ai-mcp-llm@latestImportant: Browser Installation
The crawler uses Playwright to handle dynamic content. Install Chromium once after setting up the tool:
# When running the server with uvx (recommended setup)
uvx --python 3.13 --from crawl4ai-mcp-llm@latest playwright install chromium
# From a local clone
uv run playwright install chromium🖥️ Usage
Once configured, you can use the crawler by asking your AI assistant to perform a crawl.
Usage Examples with Claude/Cline
Single Page: "Fetch https://docs.python.org/3/library/re.html (just that page) and summarize it."
Simple Crawl: "Can you crawl the site example.com and give me a summary?"
Crawl with Options: "Can you crawl https://example.com with a depth of 3 and include external links?"
Dynamic Content: "Crawl this React app and wait for the
.main-contentselector to load."Anti-bot Heuristics: "Crawl example.com with magic mode enabled."
Targeted Extraction: "Crawl the docs site but only extract content matching the
h1, p.leadCSS selector."
🧰 Tools and Resources
Tool | Purpose |
| Crawl a site (following links up to |
| Fetch exactly one page and return its Markdown, without following links or writing any file (faster for "read this page" requests) |
| Close a browser session opened with |
Resource | Content |
| JSON list of saved results (URI, name, size, date, source URL), newest first |
| Full Markdown of a saved result, e.g. |
The crawl response includes the resource URI of its result: when the returned content is truncated, the assistant can read the complete file through that resource even if it has no access to the server's file system.
Progress: crawl sends an MCP progress notification after each page. Clients that reset their timeout on progress can use a shorter timeout; otherwise keep a generous one (e.g. 600 s).
🛠️ Available Parameters (crawl tool)
The crawl tool accepts the following parameters (crawl_page accepts url, css_selector, wait_for_selector, magic, session_id, delay_before_return_html and max_content_chars):
Parameter | Type | Description | Default Value |
| string | http(s) URL to crawl (required). A bare domain such as | - |
| integer | Link depth to follow (0-5): | 2 |
| integer | Maximum number of pages to crawl (1-500). | 50 |
| boolean | Also follow links to other domains | false |
| string | CSS selector to wait for before extracting content. Useful for single-page applications. | None |
| boolean | Return the extracted content directly in the MCP response | true |
| integer | Maximum number of content characters returned in the response (1,000-500,000); the file always holds everything | 50000 |
| string | Markdown file name, always stored inside the results directory ( | automatically generated |
| boolean | Allow replacing an existing | false |
| boolean | Enable crawl4ai magic mode (anti-bot heuristics) | false |
| string | Specific CSS selector to extract only targeted elements from the page | None |
| string | Custom JavaScript code to execute before extraction (requires | None |
| string | Reuse cookies and browser state across calls with the same id | None |
| number | Delay in seconds (0-60) before extracting HTML (useful for heavy JS pages) | None |
⚙️ Configuration (environment variables)
Set these in the env section of your MCP configuration:
Variable | Description | Default |
| Directory where Markdown results are written |
|
| Allow |
|
| Allow crawling |
|
| Maximum duration of one crawl, in seconds (pages crawled so far are kept) |
|
| Maximum number of crawls running at the same time |
|
| Seconds after which an unused browser session ( |
|
| Enable crawl4ai's detailed progress logs (on stderr) |
|
| Server log level ( |
|
🔒 Security
The AI assistant decides which URLs are crawled, and crawled pages may contain prompt-injection attempts. The server therefore:
only accepts
http/httpsURLs (nofile://,raw:,data:...), and rejects hosts resolving to loopback, private or link-local addresses unlessCRAWL4AI_MCP_ALLOW_PRIVATE_NETWORKS=true;runs no custom JavaScript (including
js:wait conditions) unlessCRAWL4AI_MCP_ALLOW_JS=true;writes files only inside the results directory and never overwrites them without
overwrite=true;wraps returned page content in
<untrusted-web-content>tags so the assistant treats it as data.
Redirects and links discovered during a deep crawl are filtered on their URL only (no DNS lookup): keep private-network access disabled when the server can reach sensitive internal services.
👨💻 Development
If you want to modify the crawler or run it locally:
Clone this repository:
git clone https://github.com/laurentvv/crawl4ai-mcp-llm
cd crawl4ai-mcp-llmInstall dependencies using
uv:
uv syncTest the MCP server locally using the official MCP Inspector:
npx -y @modelcontextprotocol/inspector uv run crawl4ai-mcp-llmRun the checks (unit tests never hit the network):
uv run pytest --cov # unit tests + coverage report (80% minimum)
uv run pytest -m integration # real crawls, needs network and Chromium
uv run ruff check . && uv run ruff format --check . && uv run mypyRun the MCP server directly (for standard usage):
uv run crawl4ai-mcp-llmChanges are listed in CHANGELOG.md.
🤝 Contribution
Contributions are welcome! Feel free to open an issue or submit a pull request.
📄 License
This project is licensed under the MIT License - see the LICENSE file for details.
Available Tools
1 toolcrawlB
Crawls a website and saves its content as structured markdown to a file
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to crawl | |
| max_depth | No | Maximum crawling depth | |
| include_external | No | Whether to include external links | |
| verbose | No | Enable verbose output | |
| output_file | No | Path to output file (generated if not provided) | |
| wait_for_selector | No | CSS selector to wait for before extracting content. Useful for single-page applications. | |
| return_content | No | Whether to return the extracted content directly in the MCP response | |
| magic | No | Enable magic mode to bypass anti-bots and simulate a real browser | |
| css_selector | No | Specific CSS selector to extract only targeted elements from the page | |
| js_code | No | Custom JavaScript code to execute on the page before extraction (Requires CRAWL4AI_MCP_ALLOW_JS=true environment variable) | |
| session_id | No | Persistent session identifier to keep cookies and browser state across requests | |
| delay_before_return_html | No | Delay in seconds to wait before extracting HTML (useful for heavy JS pages) |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
With no annotations, the description carries full burden but only mentions saving to file, ignoring the return_content parameter's behavior and other complex features like magic mode and JS execution.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is very concise but underspecified for the complexity; it does not front-load key behavioral details.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the 12 parameters and lack of output schema, the description fails to explain return values, side effects, or important behavioral nuances.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema coverage is 100%, so the baseline is 3. The description adds no additional parameter meaning beyond what's in the schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the action (crawls) and the output (structured markdown to a file), leaving no ambiguity about what the tool does.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives; no context about prerequisites or appropriate scenarios.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
1 tool update
v0.1.0- First observed
crawl
TDQS
Scored across 1 tool
Only one tool exists, so there is no possibility of confusion between tools. The purpose is clear and unique.
With a single tool, naming is trivially consistent. The verb 'crawl' appropriately describes the action.
A single tool feels thin for a crawling service, which typically offers multiple options (depth, output formats, etc.). However, for a minimal markdown-only crawler, it is borderline reasonable.
The tool covers the basic crawl-and-save workflow but lacks parameters like depth, page limits, or format selection, which are common gaps for such a service.
Maintenance
Related MCP Connectors
Scrape, crawl and search the web for AI agents via MCP.
Turn any public website into an MCP server for agents to search, read and navigate.
Web scraping for AI agents. Converts URLs to clean, LLM-ready Markdown with anti-bot bypass.
Web MCP: scrape/crawl sites, web search, brand assets, app stores, YouTube, Reddit, Hacker News.
Related MCP Servers
- AlicenseAqualityAmaintenanceProvides 27 MCP-native tools for web scraping, crawling, deep research, and autonomous extraction, delivering clean Markdown and structured JSON from any website.311,766 npm2MIT
- FlicenseNot gradedqualityBmaintenanceEnables MCP-compatible AI clients to extract live web content and convert it into structured Markdown for LLM ingestion, RAG pipelines, and agentic workflows.-
- AlicenseAqualityCmaintenanceEnables local web crawling and scraping through MCP, providing tools to fetch pages as Markdown, extract structured data with CSS selectors, and capture full-page screenshots while managing a shared browser instance.31-
- AlicenseNot gradedqualityBmaintenanceEnables agents to scrape, crawl, map, search, and extract web pages as clean markdown or structured JSON directly through MCP tools.AGPL 3.0