crawl4ai-mcp
Enables AI agents in Cursor Composer to scrape single webpages and crawl websites with configurable depth and page limits.
Click on "Install Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@crawl4ai-mcpscrape the latest AI news from techcrunch.com"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
Crawl4AI MCP Server
A Model Context Protocol (MCP) server implementation that integrates Crawl4AI with Cursor AI, providing web scraping and crawling capabilities as tools for LLMs in Cursor Composer's agent mode.
System Requirements
Python 3.10 or higher installed.
Related MCP server: crawl4ai-mcp
Current Features
Single page scraping
Website crawling
Installation
Basic setup instructions also available in the Official Docs for MCP Server QuickStart.
Set up your environment
First, let's install uv and set up our Python project and environment:
MacOS/Linux:
curl -LsSf https://astral.sh/uv/install.sh | shWindows:
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"Make sure to restart your terminal afterwards to ensure that the uv command gets picked up.
After that:
Clone the repository
Install dependencies using UV:
# Navigate to the crawl4ai-mcp directory
cd crawl4ai-mcp
# Install dependencies (Only first time)
uv venv
uv sync
# Activate the venv
source .venv/bin/activate
# Run the server
python main.pyAdd to Cursor's MCP Servers or Claude's MCP Servers
You may need to put the full path to the uv executable in the command field. You can get this by running which uv on MacOS/Linux or where uv on Windows.
{
"mcpServers": {
"Crawl4AI": {
"command": "uv",
"args": [
"--directory",
"/ABSOLUTE/PATH/TO/PARENT/FOLDER/crawl4ai-mcp",
"run",
"main.py"
]
}
}
}Tools Provided
This MCP server exposes the following tools to the LLM:
scrape_webpage(url: str)Description: Scrapes the content and metadata from a single webpage using Crawl4AI.
Parameters:
url(string, required): The URL of the webpage to scrape.
Returns: A list containing a
TextContentobject with the scraped content (primarily markdown) as JSON.
crawl_website(url: str, crawl_depth: int = 1, max_pages: int = 5)Description: Crawls a website starting from the given URL up to a specified depth and page limit using Crawl4AI.
Parameters:
url(string, required): The starting URL to crawl.crawl_depth(integer, optional, default: 1): The maximum depth to crawl relative to the starting URL.max_pages(integer, optional, default: 5): The maximum number of pages to scrape during the crawl.
Returns: A list containing a
TextContentobject with a JSON array of results for the crawled pages (including URL, success status, markdown content, or error).
Available Tools
2 toolscrawl_websiteB
Crawl a website starting from the given URL up to a specified depth and page limit.
Args: url: The starting URL to crawl. crawl_depth: The maximum depth to crawl relative to the starting URL (default: 1). max_pages: The maximum number of pages to scrape during the crawl (default: 5).
Returns: List containing TextContent with a JSON array of results for crawled pages.
| Name | Required | Description | Default |
|---|---|---|---|
| crawl_depth | No | ||
| max_pages | No | ||
| url | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden. It mentions crawling behavior but lacks critical details: it doesn't specify what 'crawl' entails (e.g., following links, scraping content), whether it respects robots.txt, potential rate limits, authentication needs, or error handling. For a tool with no annotations, this leaves significant behavioral gaps.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is well-structured and appropriately sized. It begins with a clear purpose statement, followed by organized sections for Args and Returns. Each sentence serves a distinct purpose with zero wasted words, making it easy to parse.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given the tool's moderate complexity (3 parameters, no annotations, no output schema), the description covers the basic operation and parameters adequately. However, it lacks information about return format details (beyond 'List containing TextContent with a JSON array'), error conditions, and behavioral constraints that would be needed for robust use.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 0%, so the description must compensate. It provides clear semantic explanations for all three parameters: 'url' as the starting URL, 'crawl_depth' as maximum depth relative to starting URL, and 'max_pages' as maximum pages to scrape. Default values are also documented. This adds substantial value beyond the bare schema.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Crawl a website starting from the given URL up to a specified depth and page limit.' It specifies the verb ('crawl'), resource ('website'), and scope ('starting from the given URL'). However, it doesn't explicitly differentiate from the sibling tool 'scrape_webpage' beyond the crawling aspect.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
The description provides no guidance on when to use this tool versus the sibling 'scrape_webpage' or other alternatives. It mentions the action but lacks context about appropriate use cases, prerequisites, or exclusions. This leaves the agent without clear direction on tool selection.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
scrape_webpageC
Scrape content and metadata from a single webpage using Crawl4AI.
Args: url: The URL of the webpage to scrape
Returns: List containing TextContent with the result as JSON.
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
No annotations are provided, so the description carries the full burden of behavioral disclosure. It states the tool scrapes content and metadata, but lacks details on permissions, rate limits, error handling, or what 'Crawl4AI' entails (e.g., if it's a library or service). This leaves significant gaps for a web scraping tool.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is concise and well-structured with a clear purpose statement followed by Args and Returns sections. However, the 'Returns' section is vague ('List containing TextContent with the result as JSON'), which slightly reduces efficiency.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given no annotations, no output schema, and low schema coverage, the description is incomplete. It lacks details on behavioral traits, error cases, and the structure of returned data, making it inadequate for a tool that interacts with external web resources.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The description adds minimal semantics: it defines 'url' as 'The URL of the webpage to scrape'. With 0% schema description coverage and only one parameter, this provides basic meaning, but doesn't elaborate on URL format constraints or validation, leaving room for improvement.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description clearly states the tool's purpose: 'Scrape content and metadata from a single webpage using Crawl4AI.' It specifies the verb ('scrape'), resource ('content and metadata'), and scope ('single webpage'), though it doesn't explicitly differentiate from its sibling 'crawl_website' beyond implying single vs. multi-page operations.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
No guidance is provided on when to use this tool versus alternatives. The description mentions 'single webpage' and the sibling is named 'crawl_website', which might imply this is for single pages while the sibling is for entire sites, but this is not explicitly stated, leaving usage context unclear.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
TDQS
The two tools have clearly distinct purposes: crawl_website handles multi-page crawling with depth and page limits, while scrape_webpage focuses on single-page extraction. There is no overlap in functionality, making it easy for an agent to choose the right tool based on whether it needs to crawl multiple pages or scrape just one.
Both tools follow a consistent verb_noun pattern (crawl_website and scrape_webpage), using snake_case and descriptive action-object naming. This consistency makes the tool set predictable and easy to understand at a glance.
With only 2 tools, the server feels too thin for a web crawling/scraping domain. While the tools cover basic crawling and scraping, typical MCP servers in this domain would include additional utilities like filtering, parsing, or handling different content types, making this set appear incomplete and limited in scope.
The tool set is severely incomplete for web crawling and scraping. It lacks essential operations such as updating crawl parameters, deleting or managing crawl jobs, handling errors or retries, and processing extracted data (e.g., cleaning, summarizing). This forces agents into dead ends for common workflows beyond basic one-off tasks.
Maintenance
Resources
Unclaimed servers have limited discoverability.
Looking for Admin?
If you are the server author, to access and configure the admin panel.
Related MCP Connectors
MCP server for web extraction and rendering via AceDataCloud WebExtrator
Scrape, crawl and search the web for AI agents via MCP.
Zenrows MCP server — Fetch, Extract, Batch, and Browser Sessions for AI coding assistants
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceJava implementation of MCP Server for Crawl4ai4MIT
- AlicenseNot gradedqualityCmaintenanceMCP server integrating Crawl4AI for universal web crawling and data extraction. Enables AI agents to crawl, extract markdown/HTML, take screenshots, generate PDFs, and execute JavaScript on web pages.336MIT
- AlicenseNot gradedqualityDmaintenanceA lightweight MCP server that exposes Crawl4AI web scraping and crawling capabilities as tools for AI agents, enabling single-page scraping and multi-page crawling with adaptive stopping.107MIT
- AlicenseAqualityAmaintenanceMCP server for web search and crawling, integrating SearXNG metasearch and Crawl4AI for privacy-respecting search and content extraction.3161MIT
Latest Blog Posts
- Who's Calling? MCP Hosts Are an Identity Blind Spot (And the Spec Knows It)By Om-Shree-0709 on .mcpAgent IdentityOAuth 2.1
- Your AI Chatbot Just Exposed Your CEO's Salary to an InternBy Om-Shree-0709 on .Agent IdentityMCP SecurityOAuth Delegation
- Why MCP Servers Need Execution Sandboxing (And Why Your Current Stack Isn't Enough)By Om-Shree-0709 on .Agentic AiPrompt InjectionWebAssembly
MCP directory API
We provide all the information about MCP servers via our MCP API.
curl -X GET 'https://glama.ai/api/mcp/v1/servers/ritvij14/crawl4ai-mcp'
If you have feedback or need assistance with the MCP directory API, please join our Discord server