Skip to main content
Glama

Crawl4AI MCP Server

A Model Context Protocol (MCP) server implementation that integrates Crawl4AI with Cursor AI, providing web scraping and crawling capabilities as tools for LLMs in Cursor Composer's agent mode.

System Requirements

Python 3.10 or higher installed.

Related MCP server: crawl4ai-mcp

Current Features

  • Single page scraping

  • Website crawling

Installation

Basic setup instructions also available in the Official Docs for MCP Server QuickStart.

Set up your environment

First, let's install uv and set up our Python project and environment:

MacOS/Linux:

curl -LsSf https://astral.sh/uv/install.sh | sh

Windows:

powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

Make sure to restart your terminal afterwards to ensure that the uv command gets picked up.

After that:

  1. Clone the repository

  2. Install dependencies using UV:

# Navigate to the crawl4ai-mcp directory
cd crawl4ai-mcp

# Install dependencies (Only first time)
uv venv
uv sync

# Activate the venv
source .venv/bin/activate

# Run the server
python main.py
  1. Add to Cursor's MCP Servers or Claude's MCP Servers

You may need to put the full path to the uv executable in the command field. You can get this by running which uv on MacOS/Linux or where uv on Windows.

{
  "mcpServers": {
    "Crawl4AI": {
      "command": "uv",
      "args": [
        "--directory",
        "/ABSOLUTE/PATH/TO/PARENT/FOLDER/crawl4ai-mcp",
        "run",
        "main.py"
      ]
    }
  }
}

Tools Provided

This MCP server exposes the following tools to the LLM:

  1. scrape_webpage(url: str)

    • Description: Scrapes the content and metadata from a single webpage using Crawl4AI.

    • Parameters:

      • url (string, required): The URL of the webpage to scrape.

    • Returns: A list containing a TextContent object with the scraped content (primarily markdown) as JSON.

  2. crawl_website(url: str, crawl_depth: int = 1, max_pages: int = 5)

    • Description: Crawls a website starting from the given URL up to a specified depth and page limit using Crawl4AI.

    • Parameters:

      • url (string, required): The starting URL to crawl.

      • crawl_depth (integer, optional, default: 1): The maximum depth to crawl relative to the starting URL.

      • max_pages (integer, optional, default: 5): The maximum number of pages to scrape during the crawl.

    • Returns: A list containing a TextContent object with a JSON array of results for the crawled pages (including URL, success status, markdown content, or error).

Available Tools

2 tools
crawl_websiteB

Crawl a website starting from the given URL up to a specified depth and page limit.

Args: url: The starting URL to crawl. crawl_depth: The maximum depth to crawl relative to the starting URL (default: 1). max_pages: The maximum number of pages to scrape during the crawl (default: 5).

Returns: List containing TextContent with a JSON array of results for crawled pages.

ParametersJSON Schema
NameRequiredDescriptionDefault
crawl_depthNo
max_pagesNo
urlYes

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden. It mentions crawling behavior but lacks critical details: it doesn't specify what 'crawl' entails (e.g., following links, scraping content), whether it respects robots.txt, potential rate limits, authentication needs, or error handling. For a tool with no annotations, this leaves significant behavioral gaps.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and appropriately sized. It begins with a clear purpose statement, followed by organized sections for Args and Returns. Each sentence serves a distinct purpose with zero wasted words, making it easy to parse.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the tool's moderate complexity (3 parameters, no annotations, no output schema), the description covers the basic operation and parameters adequately. However, it lacks information about return format details (beyond 'List containing TextContent with a JSON array'), error conditions, and behavioral constraints that would be needed for robust use.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate. It provides clear semantic explanations for all three parameters: 'url' as the starting URL, 'crawl_depth' as maximum depth relative to starting URL, and 'max_pages' as maximum pages to scrape. Default values are also documented. This adds substantial value beyond the bare schema.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Crawl a website starting from the given URL up to a specified depth and page limit.' It specifies the verb ('crawl'), resource ('website'), and scope ('starting from the given URL'). However, it doesn't explicitly differentiate from the sibling tool 'scrape_webpage' beyond the crawling aspect.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool versus the sibling 'scrape_webpage' or other alternatives. It mentions the action but lacks context about appropriate use cases, prerequisites, or exclusions. This leaves the agent without clear direction on tool selection.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scrape_webpageC

Scrape content and metadata from a single webpage using Crawl4AI.

Args: url: The URL of the webpage to scrape

Returns: List containing TextContent with the result as JSON.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes

TDQS

C2.9/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool scrapes content and metadata, but lacks details on permissions, rate limits, error handling, or what 'Crawl4AI' entails (e.g., if it's a library or service). This leaves significant gaps for a web scraping tool.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is concise and well-structured with a clear purpose statement followed by Args and Returns sections. However, the 'Returns' section is vague ('List containing TextContent with the result as JSON'), which slightly reduces efficiency.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no annotations, no output schema, and low schema coverage, the description is incomplete. It lacks details on behavioral traits, error cases, and the structure of returned data, making it inadequate for a tool that interacts with external web resources.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description adds minimal semantics: it defines 'url' as 'The URL of the webpage to scrape'. With 0% schema description coverage and only one parameter, this provides basic meaning, but doesn't elaborate on URL format constraints or validation, leaving room for improvement.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's purpose: 'Scrape content and metadata from a single webpage using Crawl4AI.' It specifies the verb ('scrape'), resource ('content and metadata'), and scope ('single webpage'), though it doesn't explicitly differentiate from its sibling 'crawl_website' beyond implying single vs. multi-page operations.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

No guidance is provided on when to use this tool versus alternatives. The description mentions 'single webpage' and the sibling is named 'crawl_website', which might imply this is for single pages while the sibling is for entire sites, but this is not explicitly stated, leaving usage context unclear.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

TDQS

B3.1/5.0
Disambiguation5/5

The two tools have clearly distinct purposes: crawl_website handles multi-page crawling with depth and page limits, while scrape_webpage focuses on single-page extraction. There is no overlap in functionality, making it easy for an agent to choose the right tool based on whether it needs to crawl multiple pages or scrape just one.

Naming Consistency5/5

Both tools follow a consistent verb_noun pattern (crawl_website and scrape_webpage), using snake_case and descriptive action-object naming. This consistency makes the tool set predictable and easy to understand at a glance.

Tool Count2/5

With only 2 tools, the server feels too thin for a web crawling/scraping domain. While the tools cover basic crawling and scraping, typical MCP servers in this domain would include additional utilities like filtering, parsing, or handling different content types, making this set appear incomplete and limited in scope.

Completeness2/5

The tool set is severely incomplete for web crawling and scraping. It lacks essential operations such as updating crawl parameters, deleting or managing crawl jobs, handling errors or retries, and processing extracted data (e.g., cleaning, summarizing). This forces agents into dead ends for common workflows beyond basic one-off tasks.

Maintenance

ActivityInactive
ResponsivenessSyncing

Resources

Unclaimed servers have limited discoverability.

Looking for Admin?

If you are the server author, to access and configure the admin panel.

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server integrating Crawl4AI for universal web crawling and data extraction. Enables AI agents to crawl, extract markdown/HTML, take screenshots, generate PDFs, and execute JavaScript on web pages.
    33
    6
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    A lightweight MCP server that exposes Crawl4AI web scraping and crawling capabilities as tools for AI agents, enabling single-page scraping and multi-page crawling with adaptive stopping.
    107
    MIT

Latest Blog Posts

MCP directory API

We provide all the information about MCP servers via our MCP API.

curl -X GET 'https://glama.ai/api/mcp/v1/servers/ritvij14/crawl4ai-mcp'

If you have feedback or need assistance with the MCP directory API, please join our Discord server