Skip to main content
Glama
laurentvv

Crawl4AI MCP

by laurentvv

Web Crawler MCP

English 中文 हिंदी Español Français العربية বাংলা Русский Português Bahasa Indonesia

Python License

A powerful web crawling tool that integrates with AI assistants via the MCP (Model Context Protocol). This project allows AI assistants to crawl websites, extract dynamic content, navigate through links, and save structured Markdown files directly.

📋 Features

  • Native integration with AI assistants via MCP

  • Return scraped Markdown content directly to the AI

  • Extracts and surfaces internal/external links for AI navigation

  • Website crawling with configurable depth and page limit

  • Detailed crawl result statistics, including the list of skipped pages

  • Live progress notifications during long crawls

  • Saved results exposed as MCP resources (crawl://results/...)

  • Error and not found page handling

  • Secure by default: only public http(s) URLs, JavaScript disabled unless you opt in

  • Advanced Scraping Capabilities:

    • Magic Mode: crawl4ai heuristics that simulate a real browser and help with some anti-bot protections (not a guaranteed bypass)

    • Targeted Extraction: Fetch only what you need using CSS selectors

    • Custom JavaScript (opt-in): Execute code before extraction (clicks, scrolls, form fills)

    • Persistent Sessions: One browser is shared across calls, so a session_id keeps cookies and state until it is closed (close_session) or stays idle too long

    • SPA Support: Wait for dynamic CSS selectors or set explicit pre-extraction delays

Related MCP server: Scrapling MCP Server

🚀 MCP Configuration

The simplest and recommended way to use this tool is via uvx. The @latest suffix makes uvx check PyPI at each start and run the newest published version (without it, uvx keeps reusing the version it cached the first time).

Prerequisites

  • uv installed on your system.

Setup for AI Assistants (e.g., Claude Desktop, Cline)

Add the following to your AI Assistant's MCP configuration file (e.g., cline_mcp_settings.json or claude_desktop_config.json):

Note: Python 3.12 or 3.13 is required (crawl4ai does not support 3.14 yet). Specifying --python 3.13 is recommended, especially on Windows, to avoid compilation issues with certain dependencies.

From PyPI (recommended):

{
  "mcpServers": {
    "crawl": {
      "command": "uvx",
      "args": [
        "--python",
        "3.13",
        "crawl4ai-mcp-llm@latest"
      ],
      "disabled": false,
      "autoApprove": [],
      "timeout": 600
    }
  }
}

From GitHub (latest unreleased):

{
  "mcpServers": {
    "crawl": {
      "command": "uvx",
      "args": [
        "--python",
        "3.13",
        "--from",
        "git+https://github.com/laurentvv/crawl4ai-mcp-llm",
        "crawl4ai-mcp-llm"
      ],
      "disabled": false,
      "autoApprove": [],
      "timeout": 600
    }
  }
}

Claude Code:

claude mcp add crawl -s user -- uvx --python 3.13 crawl4ai-mcp-llm@latest

Important: Browser Installation

The crawler uses Playwright to handle dynamic content. Install Chromium once after setting up the tool:

# When running the server with uvx (recommended setup)
uvx --python 3.13 --from crawl4ai-mcp-llm@latest playwright install chromium

# From a local clone
uv run playwright install chromium

🖥️ Usage

Once configured, you can use the crawler by asking your AI assistant to perform a crawl.

Usage Examples with Claude/Cline

  • Single Page: "Fetch https://docs.python.org/3/library/re.html (just that page) and summarize it."

  • Simple Crawl: "Can you crawl the site example.com and give me a summary?"

  • Crawl with Options: "Can you crawl https://example.com with a depth of 3 and include external links?"

  • Dynamic Content: "Crawl this React app and wait for the .main-content selector to load."

  • Anti-bot Heuristics: "Crawl example.com with magic mode enabled."

  • Targeted Extraction: "Crawl the docs site but only extract content matching the h1, p.lead CSS selector."

🧰 Tools and Resources

Tool

Purpose

crawl

Crawl a site (following links up to max_depth/max_pages), save the result as Markdown and return a summary with the content

crawl_page

Fetch exactly one page and return its Markdown, without following links or writing any file (faster for "read this page" requests)

close_session

Close a browser session opened with session_id (idle sessions are also closed automatically after CRAWL4AI_MCP_SESSION_TTL)

Resource

Content

crawl://results

JSON list of saved results (URI, name, size, date, source URL), newest first

crawl://results/{path}

Full Markdown of a saved result, e.g. crawl://results/crawl_example_com_20260101_120000_ab12cd.md

The crawl response includes the resource URI of its result: when the returned content is truncated, the assistant can read the complete file through that resource even if it has no access to the server's file system.

Progress: crawl sends an MCP progress notification after each page. Clients that reset their timeout on progress can use a shorter timeout; otherwise keep a generous one (e.g. 600 s).

🛠️ Available Parameters (crawl tool)

The crawl tool accepts the following parameters (crawl_page accepts url, css_selector, wait_for_selector, magic, session_id, delay_before_return_html and max_content_chars):

Parameter

Type

Description

Default Value

url

string

http(s) URL to crawl (required). A bare domain such as example.com is treated as https://example.com.

-

max_depth

integer

Link depth to follow (0-5): 0 = start page only, 1 = the start page and its links, and so on

2

max_pages

integer

Maximum number of pages to crawl (1-500). 1 fetches exactly one page.

50

include_external

boolean

Also follow links to other domains

false

wait_for_selector

string

CSS selector to wait for before extracting content. Useful for single-page applications.

None

return_content

boolean

Return the extracted content directly in the MCP response

true

max_content_chars

integer

Maximum number of content characters returned in the response (1,000-500,000); the file always holds everything

50000

output_file

string

Markdown file name, always stored inside the results directory (.md is added if missing)

automatically generated

overwrite

boolean

Allow replacing an existing output_file

false

magic

boolean

Enable crawl4ai magic mode (anti-bot heuristics)

false

css_selector

string

Specific CSS selector to extract only targeted elements from the page

None

js_code

string

Custom JavaScript code to execute before extraction (requires CRAWL4AI_MCP_ALLOW_JS=true)

None

session_id

string

Reuse cookies and browser state across calls with the same id

None

delay_before_return_html

number

Delay in seconds (0-60) before extracting HTML (useful for heavy JS pages)

None

⚙️ Configuration (environment variables)

Set these in the env section of your MCP configuration:

Variable

Description

Default

CRAWL4AI_RESULTS_DIR

Directory where Markdown results are written

~/.crawl4ai_mcp_llm/results

CRAWL4AI_MCP_ALLOW_JS

Allow js_code and JavaScript wait conditions (js:...)

false

CRAWL4AI_MCP_ALLOW_PRIVATE_NETWORKS

Allow crawling localhost and private/link-local addresses

false

CRAWL4AI_MCP_CRAWL_TIMEOUT

Maximum duration of one crawl, in seconds (pages crawled so far are kept)

300

CRAWL4AI_MCP_MAX_CONCURRENT_CRAWLS

Maximum number of crawls running at the same time

2

CRAWL4AI_MCP_SESSION_TTL

Seconds after which an unused browser session (session_id) is closed

1800

CRAWL4AI_MCP_VERBOSE

Enable crawl4ai's detailed progress logs (on stderr)

false

CRAWL4AI_MCP_LOG_LEVEL

Server log level (DEBUG, INFO, WARNING, ...)

INFO

🔒 Security

The AI assistant decides which URLs are crawled, and crawled pages may contain prompt-injection attempts. The server therefore:

  • only accepts http/https URLs (no file://, raw:, data: ...), and rejects hosts resolving to loopback, private or link-local addresses unless CRAWL4AI_MCP_ALLOW_PRIVATE_NETWORKS=true;

  • runs no custom JavaScript (including js: wait conditions) unless CRAWL4AI_MCP_ALLOW_JS=true;

  • writes files only inside the results directory and never overwrites them without overwrite=true;

  • wraps returned page content in <untrusted-web-content> tags so the assistant treats it as data.

Redirects and links discovered during a deep crawl are filtered on their URL only (no DNS lookup): keep private-network access disabled when the server can reach sensitive internal services.

👨‍💻 Development

If you want to modify the crawler or run it locally:

  1. Clone this repository:

git clone https://github.com/laurentvv/crawl4ai-mcp-llm
cd crawl4ai-mcp-llm
  1. Install dependencies using uv:

uv sync
  1. Test the MCP server locally using the official MCP Inspector:

npx -y @modelcontextprotocol/inspector uv run crawl4ai-mcp-llm
  1. Run the checks (unit tests never hit the network):

uv run pytest --cov                # unit tests + coverage report (80% minimum)
uv run pytest -m integration       # real crawls, needs network and Chromium
uv run ruff check . && uv run ruff format --check . && uv run mypy
  1. Run the MCP server directly (for standard usage):

uv run crawl4ai-mcp-llm

Changes are listed in CHANGELOG.md.

🤝 Contribution

Contributions are welcome! Feel free to open an issue or submit a pull request.

📄 License

This project is licensed under the MIT License - see the LICENSE file for details.

Glama MCP

Available Tools

1 tool
crawlB

Crawls a website and saves its content as structured markdown to a file

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesURL to crawl
max_depthNoMaximum crawling depth
include_externalNoWhether to include external links
verboseNoEnable verbose output
output_fileNoPath to output file (generated if not provided)
wait_for_selectorNoCSS selector to wait for before extracting content. Useful for single-page applications.
return_contentNoWhether to return the extracted content directly in the MCP response
magicNoEnable magic mode to bypass anti-bots and simulate a real browser
css_selectorNoSpecific CSS selector to extract only targeted elements from the page
js_codeNoCustom JavaScript code to execute on the page before extraction (Requires CRAWL4AI_MCP_ALLOW_JS=true environment variable)
session_idNoPersistent session identifier to keep cookies and browser state across requests
delay_before_return_htmlNoDelay in seconds to wait before extracting HTML (useful for heavy JS pages)

TDQS

B3/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries full burden but only mentions saving to file, ignoring the return_content parameter's behavior and other complex features like magic mode and JS execution.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is very concise but underspecified for the complexity; it does not front-load key behavioral details.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the 12 parameters and lack of output schema, the description fails to explain return values, side effects, or important behavioral nuances.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3. The description adds no additional parameter meaning beyond what's in the schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the action (crawls) and the output (structured markdown to a file), leaving no ambiguity about what the tool does.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives; no context about prerequisites or appropriate scenarios.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 1 tool updatev0.1.0
    • First observedcrawl

TDQS

B3.3/5.0

Scored across 1 tool

Disambiguation5/5

Only one tool exists, so there is no possibility of confusion between tools. The purpose is clear and unique.

Naming Consistency5/5

With a single tool, naming is trivially consistent. The verb 'crawl' appropriately describes the action.

Tool Count3/5

A single tool feels thin for a crawling service, which typically offers multiple options (depth, output formats, etc.). However, for a minimal markdown-only crawler, it is borderline reasonable.

Completeness3/5

The tool covers the basic crawl-and-save workflow but lacks parameters like depth, page limits, or format selection, which are common gaps for such a service.

Maintenance

ActivityActive
ResponsivenessResponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    A
    quality
    C
    maintenance
    Enables local web crawling and scraping through MCP, providing tools to fetch pages as Markdown, extract structured data with CSS selectors, and capture full-page screenshots while managing a shared browser instance.
    3
    1
    -
  • A
    license
    Not graded
    quality
    B
    maintenance
    Enables agents to scrape, crawl, map, search, and extract web pages as clean markdown or structured JSON directly through MCP tools.
    AGPL 3.0