Skip to main content
Glama
bidouilles

cloudflare-crawl-mcp

by bidouilles

cloudflare-crawl-mcp

MCP server for Cloudflare Browser Rendering Crawl API. Fetches and crawls web pages, returning clean Markdown optimized for LLM consumption.

Built with FastMCP and uv.

Tools

Tool

Description

scrape_url

Fetch a single page as Markdown. The primary tool — use this when you know the URL.

map_url

Discover URLs on a site without fetching content. Use to find the right page first.

crawl_url

Crawl multiple pages and return all content as Markdown.

Typical workflow: map_url to find pages → scrape_url to read the right one.

Related MCP server: @hauntapi/mcp-server

Prerequisites

  1. A Cloudflare account with Browser Rendering enabled

  2. An API token with Account > Browser Rendering > Edit permission (create one here)

  3. uv installed

Configuration

Variable

Required

Description

CF_API_TOKEN

Yes

Cloudflare API token

CF_ACCOUNT_ID

Yes

Cloudflare Account ID

CF_RATE_LIMIT

No

API requests per minute (default: 6 for Free, set to 600 for Paid)

Setup

Claude Code

claude mcp add cloudflare-crawl \
  -e CF_API_TOKEN=your_api_token \
  -e CF_ACCOUNT_ID=your_account_id \
  -- uv run --directory /path/to/cloudflare-crawl-mcp python server.py

Codex

codex mcp add cloudflare-crawl \
  -- env CF_API_TOKEN="your_api_token" CF_ACCOUNT_ID="your_account_id" \
  uv run --directory /path/to/cloudflare-crawl-mcp python server.py

Claude Desktop

Add to ~/Library/Application Support/Claude/claude_desktop_config.json (macOS) or %APPDATA%\Claude\claude_desktop_config.json (Windows):

{
  "mcpServers": {
    "cloudflare-crawl": {
      "command": "uv",
      "args": ["run", "--directory", "/path/to/cloudflare-crawl-mcp", "python", "server.py"],
      "env": {
        "CF_API_TOKEN": "your_api_token",
        "CF_ACCOUNT_ID": "your_account_id"
      }
    }
  }
}

Cursor

Add to .cursor/mcp.json in your project or ~/.cursor/mcp.json globally:

{
  "mcpServers": {
    "cloudflare-crawl": {
      "command": "uv",
      "args": ["run", "--directory", "/path/to/cloudflare-crawl-mcp", "python", "server.py"],
      "env": {
        "CF_API_TOKEN": "your_api_token",
        "CF_ACCOUNT_ID": "your_account_id"
      }
    }
  }
}

Tool Reference

scrape_url

Fetch a single page and return its content as Markdown.

url (string, required)    — The URL to fetch
render (boolean, optional) — Render JavaScript with headless browser (default: true, set false for static pages)

map_url

Discover URLs on a website without fetching full content.

url (string, required)                  — Starting URL
limit (number, optional)                — Max URLs to discover (default: 50)
depth (number, optional)                — Link depth to follow (default: 2)
include_subdomains (boolean, optional)  — Follow subdomain links (default: false)
include_external_links (boolean, optional) — Follow external links (default: false)
include_patterns (string[], optional)   — Only visit matching URLs (e.g. "https://example.com/docs/**")
exclude_patterns (string[], optional)   — Skip matching URLs

crawl_url

Crawl multiple pages and return all content as Markdown.

url (string, required)                  — Starting URL
limit (number, optional)                — Max pages to crawl (default: 10)
depth (number, optional)                — Link depth (default: 1)
include_subdomains (boolean, optional)  — Follow subdomain links (default: false)
include_external_links (boolean, optional) — Follow external links (default: false)
include_patterns (string[], optional)   — Only visit matching URLs
exclude_patterns (string[], optional)   — Skip matching URLs
render (boolean, optional)              — Render JavaScript (default: true)

Cloudflare Plan Limits

Free

Paid

Browser time

10 min/day

10 hrs/month

API rate limit

6 req/min

600 req/min

Concurrent browsers

3

10

Max pages per job

100,000

100,000

Max job duration

7 days

7 days

Results available

14 days

14 days

render: false crawls run on Workers instead of a headless browser and do not consume browser time.

License

MIT

Available Tools

3 tools
crawl_urlA

Crawl multiple pages starting from a URL and return all content as Markdown.

Best for: Fetching entire documentation sections, blog archives, or multiple related pages at once.

Not recommended for: Single pages (use scrape_url — it's faster). Large sites without filters (responses can be very large and exceed token limits).

Tip: Use include_patterns to scope the crawl (e.g. "https://example.com/docs/**").

Args: url: The starting URL to crawl. limit: Maximum number of pages to crawl (default: 10, max: 100000). depth: Maximum link depth from the starting URL (default: 1). include_subdomains: If true, follows links to subdomains. include_external_links: If true, follows links to external domains. include_patterns: Only visit URLs matching these wildcard patterns. exclude_patterns: Skip URLs matching these wildcard patterns. render: If true (default), renders JavaScript. Set false for faster static fetch.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
depthNo
limitNo
renderNo
exclude_patternsNo
include_patternsNo
include_subdomainsNo
include_external_linksNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden and does so thoroughly. It explains default behavior (render JavaScript, depth 1, limit 10), optional behaviors (include_subdomains, include_external_links, patterns), and warns about large responses exceeding token limits. It also states the output format (Markdown) and the trade-off of disabling render.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured and front-loaded: it states the core function first, then best-use cases, exclusions, a practical tip, and a clear Args list. Every sentence serves a purpose; length is justified by the tool's complexity and the absence of schema-level descriptions.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given the complex 8-parameter tool, no annotations, and an output schema, the description still manages to cover all needed context: purpose, use cases, exclusions, parameter semantics, default behaviors, and a warning about potential token limits. It is complete enough for an agent to select and invoke the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema has 0% description coverage, so the description must compensate—and it does. The 'Args' section explains every parameter's meaning, defaults, and constraints (e.g., 'limit: Maximum number of pages to crawl (default: 10, max: 100000)'). This adds significant semantic value beyond the bare schema types.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's function: 'Crawl multiple pages starting from a URL and return all content as Markdown.' It uses a specific verb ('crawl'), names the resource ('a URL'), and explicitly differentiates from siblings by noting 'Single pages (use scrape_url — it's faster)'.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit when-to-use guidance: 'Best for' fetching entire documentation sections, blog archives, or multiple related pages. It also gives clear when-not-to-use guidance and names the alternative tool (scrape_url) for single pages. The tip about using include_patterns to scope crawls adds practical usage context.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

map_urlA

Discover URLs on a website without fetching full page content.

Returns a list of URLs found by crawling from the starting URL. Use this to find the right page before scraping it.

Best for: Finding documentation pages, locating specific content on a site, understanding site structure before scraping.

Typical workflow: map_url to find URLs -> scrape_url on the right page.

Args: url: The starting URL to discover links from. limit: Maximum number of URLs to discover (default: 50, max: 100000). depth: How many links deep to follow (default: 2). include_subdomains: If true, follows links to subdomains. include_external_links: If true, follows links to external domains. include_patterns: Only visit URLs matching these wildcard patterns (e.g. "https://example.com/docs/**"). exclude_patterns: Skip URLs matching these wildcard patterns.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
depthNo
limitNo
exclude_patternsNo
include_patternsNo
include_subdomainsNo
include_external_linksNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description discloses the key behavior of not fetching full page content, that it returns a list of URLs, and that it crawls with configurable depth/limit and include/exclude patterns. It clearly states it 'discovers' URLs, implying a read-only operation.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a clear intro, best-for list, typical workflow, and concise args explanations. Each sentence adds value; no wordiness.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

With an output schema present, the description doesn't need to explain return values. It covers purpose, usage context, parameter semantics, and behavior comprehensively, making it complete for an agent to select and invoke the tool.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The description provides a detailed Args section that explains every parameter, including the meaning of limit, depth, include_subdomains, include_external_links, include_patterns, and exclude_patterns, with an example wildcard for include_patterns. This fully compensates for the 0% schema coverage.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states 'Discover URLs on a website without fetching full page content' and 'Returns a list of URLs found by crawling from the starting URL.' This clearly identifies the resource (website URLs) and specific verb (discover), and distinguishes it from sibling tools like scrape_url, which fetches content.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

It explicitly says 'Use this to find the right page before scraping it' and includes a 'Typical workflow: map_url to find URLs -> scrape_url on the right page.' It also lists best-for use cases, providing clear guidance on when to use and an alternative.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

scrape_urlA

Fetch a single web page and return its content as Markdown.

This is the primary tool for reading web content. Use this when you know which URL contains the information you need.

Best for: Reading documentation pages, articles, blog posts, API references, or any single page where you know the URL.

Not recommended for: When you don't know which page has the info (use map_url first), or when you need content from many pages (use crawl_url).

Args: url: The URL to fetch. render: If true (default), renders JavaScript with a headless browser. Set to false for faster fetch of static pages.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYes
renderNo

Output Schema

ParametersJSON Schema
NameRequiredDescription
resultYes

TDQS

A4.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations provided, the description carries the full burden. It discloses the JavaScript rendering behavior via the render flag and notes the performance tradeoff, adding behavioral detail beyond the schema.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-structured with a clear purpose statement, best-for/not-recommended sections, and an Args list. Every sentence adds useful context without redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool's behavior, use cases, and parameter options are fully described. The output is described as Markdown content, and the presence of an output schema (not shown here) would further define the return value, so the description is sufficiently complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has zero description coverage, but the 'Args' section in the description explains both parameters. It defines 'url' as the URL to fetch and explains 'render' including its default value and effect on performance.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states the tool's action: 'Fetch a single web page and return its content as Markdown.' It identifies the specific resource (web page) and distinguishes itself from siblings by noting when to use it versus map_url and crawl_url.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides explicit usage guidance: 'Use this when you know which URL contains the information you need.' It also gives clear exclusions and alternatives, saying to use map_url when the URL is unknown and crawl_url for multiple pages.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv1.0.0
    • First observedcrawl_url
    • First observedmap_url
    • First observedscrape_url

TDQS

A4.9/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a clear, distinct purpose: scrape_url fetches one page, map_url discovers URLs without fetching content, and crawl_url fetches multiple pages. The descriptions explicitly state when to use each, so there is no ambiguity.

Naming Consistency5/5

All tool names follow a consistent verb_noun pattern with underscores: scrape_url, map_url, crawl_url. The style is uniform and predictable, making it easy to infer functionality.

Tool Count5/5

Three tools is well-scoped for a crawling server. Each tool covers a distinct core operation—single fetch, URL discovery, and bulk crawl—without unnecessary redundancy or bloat.

Completeness4/5

The tools cover the essential crawl lifecycle: discover URLs, scrape a single page, and crawl multiple pages. Minor gaps exist (e.g., no sitemap parsing or headless options for custom headers), but the core workflows are complete.

Maintenance

ActivityInactive
ResponsivenessNo issues

Related MCP Connectors

Related MCP Servers