crawl4tools
This server is an MCP server that fetches web pages and returns or saves their content in multiple formats.
fetchreturns page content directly as Markdown, HTML, or a full-page PNG screenshot; PDFs are transcribed to Markdown and images are returned as images.downloadsaves pages to files on the server in Markdown, HTML, PDF, screenshot, MHTML, or raw source format, returning the saved paths.Supports multiple URLs per call (up to the server limit), fetched concurrently and deduplicated.
Failed URLs are reported individually; the call fails only if every URL fails.
Markdown-specific options: keep only the main content (
fit), add numbered link references (citations), and drop links or images.Per-call timeout can be set with
timeout_s.downloadcan save into a chosen subdirectory of the server's download directory.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@crawl4toolsGet the full content of https://example.com as Markdown"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
crawl4tools
What is crawl4tools
crawl4tools is a web crawler built on top of the crawl4ai library. It is designed to serve three purposes:
An HTTP server,
crawl4server, that works directly as an Open WebUI external web loader, returning the body of a URL as Markdown. On the Open WebUI side, this only requires setting Web Loader Engine toexternaland the External Web Loader URL in the admin UI (or the equivalentWEB_LOADER_ENGINE=externalandEXTERNAL_WEB_LOADER_URLenvironment variables).An MCP server that lets AI agents such as Claude Code fetch the full content of a web page, rather than a summary.
A local CLI that downloads a given URL (or multiple URLs at once) as Markdown or other formats.
Related MCP server: cleanfetch
Status
Alpha. This release (0.4.0) provides the local CLI crawl4cli, the MCP server crawl4mcp, and the combined Open WebUI web loader + MCP server crawl4server, with a Dockerfile and compose file. All three commands can show their messages in English or Japanese.
Features
Available now (CLI):
Download one or more URLs as Markdown, HTML, PDF, screenshot (PNG), MHTML, or the raw source
PDFs are transcribed to Markdown; images and other non-HTML files are saved as they are
Clear error messages for HTTP errors, unknown hosts, refused connections, timeouts, and a missing browser
Downloading through an HTTP/HTTPS/SOCKS5 proxy, with a single retry over a direct connection when the proxy itself fails
Messages in English or Japanese (see Language of messages)
Available now (MCP server):
Two tools:
fetchreturns pages directly to the client as Markdown (default), HTML, or a PNG screenshot (images come back as images, PDFs are transcribed to Markdown);downloadsaves pages in any format (Markdown, HTML, PDF, screenshot, MHTML, or raw) into a directory on the server and returns the pathsSeveral URLs per call (default limit 20), fetched concurrently under a server-wide limit (default 3), sharing one headless browser
stdio (default) and Streamable HTTP transports
Proxy support with a direct-connection fallback, configured on the server side
Settings via command-line options,
CRAWL4MCP_*environment variables, or a YAML/JSON config fileMessages in English or Japanese (see Language of messages)
Available now (Open WebUI web loader, crawl4server):
POST /crawlaccepts{"urls": [...]}and returns Markdown plus metadata (source, URL, title, status code, content type) for each page Open WebUI can use; invalid or failing URLs are simply left out of the response instead of failing the whole batchServes the same MCP endpoint as
crawl4mcpon a second port, sharing one headless browser and one concurrency limit with the web loaderGET /healthfor liveness checks, and an optional bearer API key on the web loader endpointSettings via command-line options,
CRAWL4SERVER_*environment variables, or a YAML/JSON config fileDockerfile and compose file for running the server in a container
Messages in English or Japanese (see Language of messages)
Planned:
Prebuilt Docker image on a registry
Authentication for the MCP endpoint
Fallback to a direct connection when SSL bumping by the proxy breaks the result
Installation
The CLI and the MCP server need Python 3.11 or newer and uv.
uv tool install --with-executables-from playwright git+https://github.com/xhighhongo41/crawl4tools
playwright install chromium # downloads the headless browser (once)This installs crawl4cli, crawl4mcp, and crawl4server.
The browser is stored in Playwright's cache directory (for example ~/Library/Caches/ms-playwright on macOS). crawl4ai also creates a ~/.crawl4ai directory for its own data.
To run crawl4server in a container instead, see Docker below.
Usage
crawl4cli https://example.com/ # Markdown to stdout
crawl4cli -o page.md https://example.com/ # save to a file
crawl4cli -d out/ URL1 URL2 URL3 # several URLs into a directory
crawl4cli -f screenshot https://example.com/ # saves example.com.png
crawl4cli --proxy http://proxy.local:8080 URL # through a proxyOption | Meaning |
|
|
| Save a single URL to FILE instead of stdout |
| Directory for several URLs and binary formats (default: current directory) |
|
|
| Do not retry over a direct connection when the proxy fails |
| URLs fetched at once (default: 3) |
| Page load timeout per URL (default: 60) |
| Keep only the main content (drops menus, footers, and the like); falls back to the full page if nothing is left |
| Turn links into numbered references listed at the end |
| Drop links or image references from the Markdown |
| Less or more output on stderr |
| Language of messages (default: follows the OS locale; see Language of messages) |
Only the document goes to stdout; notes, errors, and the summary go to stderr. File names are derived from the URL (https://example.com/a/b → example.com_a_b.md). The exit code is 0 when every URL succeeded, 1 when any failed, and 2 for invalid arguments. Every option can also be set with an environment variable named CRAWL4CLI_<OPTION>, for example CRAWL4CLI_PROXY. The standard HTTP_PROXY/HTTPS_PROXY variables are not used.
Open WebUI web loader (crawl4server)
crawl4server runs a single process that serves both an Open WebUI external web loader and an MCP server, on two separate ports, sharing one headless browser and one concurrency limit (-j) between them.
Running
crawl4serverBy default this listens on 127.0.0.1:8766 for the web loader (POST /crawl, GET /health) and 127.0.0.1:8765 for MCP (/mcp).
Connecting Open WebUI
In Open WebUI (0.11.x or later), under Admin Settings > Web Search, set:
Web Loader Engine:
externalExternal Web Loader URL:
http://<host>:8766/crawlExternal Web Loader API Key: the value passed to
--loader-api-key(leave empty if none)
The equivalent environment variables on the Open WebUI side are WEB_LOADER_ENGINE=external, EXTERNAL_WEB_LOADER_URL, and EXTERNAL_WEB_LOADER_API_KEY.
Open WebUI posts {"urls": [...]} to that URL and gets back a JSON array of {"page_content": <Markdown>, "metadata": {"source", "url", "title", "status_code", "content_type"}} documents. URLs that are invalid or fail are simply left out of the response (and logged to the server's stderr), so one bad URL never drops the whole batch. Open WebUI does not time out these requests itself, so tune --timeout and -j/--concurrency if searches feel slow.
Options
Option | Meaning |
| Host to listen on, both ports (default: |
| Web loader port (default: |
| HTTP path of the web loader endpoint (default: |
| Require |
| Keep only the main content of each page (default: off, full-page Markdown) |
| MCP Streamable HTTP port (default: |
| HTTP path of the MCP endpoint (default: |
|
|
| Do not retry over a direct connection when the proxy fails |
| Default per-URL timeout, for web loader requests and MCP tool calls that omit |
| Maximum URLs fetched at once across both ports (default: 3) |
| Maximum URLs accepted per web loader request or MCP tool call; Open WebUI sends up to 20, so keep this at 20 or more (default: 20) |
| Root directory the MCP |
| YAML or JSON config file (see Configuration below) |
| Verbose logging on stderr |
| Language of messages the server produces while running (default: English; see Language of messages) |
Configuration
Settings are resolved in this order: command-line options > CRAWL4SERVER_* environment variables (for example CRAWL4SERVER_LOADER_API_KEY) > config file (--config or CRAWL4SERVER_CONFIG) > built-in defaults. Config file keys are the same option names in snake_case:
host: 0.0.0.0
loader_port: 8766
loader_api_key: change-me
mcp_port: 8765
max_urls: 20
concurrency: 3
lang: jaDocker
The repository ships a Dockerfile and compose.yaml (no image is published to a registry yet, so build it locally):
git clone https://github.com/xhighhongo41/crawl4tools
cd crawl4tools
mkdir -p downloads
docker compose up -d --build
curl http://localhost:8766/healthSet CRAWL4SERVER_LOADER_API_KEY in compose.yaml's environment section. The container runs as uid 1000, so downloads/ must be writable by it; shm_size: 1gb is set for Chromium. When Open WebUI runs in the same compose project, point it at http://crawl4tools:8766/crawl instead of localhost.
Security
The web loader checks the API key only when --loader-api-key is set; the MCP port has no authentication at all. When listening on a non-loopback host (as in the Docker setup), set a loader API key and do not expose the MCP port beyond a trusted network — crawl4server prints a warning to stderr at startup in that case.
MCP server
crawl4mcp exposes the same fetching engine as an MCP server, with two tools: fetch (return content directly) and download (save it to files). crawl4server (above) serves this same MCP endpoint alongside the Open WebUI web loader, from one process.
Connecting a client
For Claude Code, over stdio (the default transport):
claude mcp add --transport stdio crawl4tools -- crawl4mcpOr, over Streamable HTTP: start the server, then point Claude Code at it:
crawl4mcp --transport http
claude mcp add --transport http crawl4tools http://127.0.0.1:8765/mcpFor Claude Desktop, add an entry to claude_desktop_config.json:
{
"mcpServers": {
"crawl4tools": {
"command": "crawl4mcp",
"args": ["--download-dir", "/path/to/downloads"]
}
}
}If Claude Desktop cannot find crawl4mcp (it does not always see your shell's PATH), replace "command": "crawl4mcp" with the full path shown by which crawl4mcp.
Other clients that support the Streamable HTTP transport (for example Open WebUI's MCP support) can connect to http://<host>:<port>/mcp once the server is running with --transport http.
Tools
Tool | Parameters |
|
|
|
|
urls: one or morehttp/httpsURLs, up to the server's per-call limit (--max-urls, default 20); duplicate URLs are fetched once.format: forfetch, the page comes back directly asmarkdown(default),html, or ascreenshot(PNG); image URLs are returned as images, and PDFs are transcribed to Markdown. Fordownload, the file is saved asmarkdown(default),html,pdf,screenshot,mhtml, orraw.fit,citations,ignore_links,ignore_images: the same content-shaping options as the CLI's--fit,--citations,--no-links, and--no-images(Markdown only).timeout_s: page load timeout for this call, in seconds; defaults to the server's--timeout.directory(downloadonly): subdirectory of the server's download directory (--download-dir) to save into; it cannot resolve outside that directory. Defaults to the download directory itself.
When a call covers several URLs, each result starts with a <!-- crawl4tools: url=... status=... --> line; a URL that failed is reported as an error: ... line instead, and the call only fails when every URL fails. A proxy fallback or other remark about a result appears as a <!-- note: ... --> line. download overwrites files that already exist at the destination.
Options
Option | Meaning |
|
|
| Host to listen on (http transport only; default: |
| Port to listen on (http transport only; default: |
| HTTP path for the MCP endpoint (http transport only; default: |
|
|
| Do not retry over a direct connection when the proxy fails |
| Default per-URL timeout, used when a tool call omits |
| Maximum URLs fetched at once across every tool call (default: 3) |
| Maximum URLs accepted in a single tool call (default: 20) |
| Root directory the |
| YAML or JSON config file (see Configuration below) |
| Verbose logging on stderr |
| Language of messages the server produces while running (default: English; see Language of messages) |
Configuration
Settings are resolved in this order: command-line options > CRAWL4MCP_* environment variables > config file > built-in defaults.
Every option can be set with an environment variable named CRAWL4MCP_ followed by the option's name in upper case with dashes replaced by underscores — for example CRAWL4MCP_MAX_URLS for --max-urls, or CRAWL4MCP_CONCURRENCY for --concurrency. CRAWL4MCP_CONFIG points to the config file itself, same as --config.
The config file is YAML (JSON also works, since JSON is valid YAML), with the same option names as keys, using underscores instead of dashes; unknown keys are an error:
transport: http
host: 127.0.0.1
port: 8765
max_urls: 50
concurrency: 5
download_dir: ./downloads
proxy: http://proxy.local:8080
lang: jaRelative paths in the config file (such as download_dir) are resolved from the current directory the server is started in.
Security
The Streamable HTTP transport has no authentication. By default the server listens on 127.0.0.1 only; if you bind it to another host, anyone who can reach that port can use the server, and crawl4mcp prints a warning to stderr when it starts. In stdio mode, stdout is reserved for the MCP protocol — all logging goes to stderr.
Language of messages
All three commands can show their messages in English (en) or Japanese (ja).
Choose the language with the --lang option, an environment variable (CRAWL4CLI_LANG, CRAWL4MCP_LANG, CRAWL4SERVER_LANG), or, for the two servers, the lang key in the config file. When more than one is set, the option wins, then the environment variable, then the config file.
Defaults: crawl4cli follows the OS locale (LANGUAGE, LC_ALL, LC_MESSAGES, then LANG, whichever is set first) and uses Japanese when it starts with ja (for example LANG=ja_JP.UTF-8), otherwise English. crawl4mcp and crawl4server always default to English, regardless of the locale, so containers and MCP clients get stable output.
crawl4cli --lang ja https://example.com/
CRAWL4SERVER_LANG=ja crawl4serverThis covers --help, error messages, notes and progress lines on stderr, the MCP tools' descriptions and result text, the servers' startup/warning lines, and the web loader's JSON error responses. --help always follows --lang/the environment variable (and, for crawl4cli, the locale) — the config file's lang only applies to messages produced while the server is running.
Always in English, unchanged by --lang: log output; the error:, note:, saved:, done:, failed: line prefixes; JSON keys; the <!-- crawl4tools: url=... status=... --> header lines; GET /health; --version; and anything printed by click or other libraries (for example Usage: or Error: Invalid value ...).
Each process uses one language for its whole run; there is no per-request language.
Acknowledgements
This product includes software developed by UncleCode (https://x.com/unclecode) as part of the Crawl4AI project (https://github.com/unclecode/crawl4ai). Crawl4AI is licensed under the Apache License 2.0.
License
This project is licensed under the Apache License 2.0. See the LICENSE file for details.
Available Tools
2 toolsdownloadDownload web pages to filesAIdempotent
Fetch one or more URLs and save each result as a file in a directory on the server, returning the saved paths. Supports Markdown (default), HTML, PDF, PNG screenshot, MHTML, and the raw source. PDFs are transcribed to Markdown unless format is 'raw'; non-web content (images, archives, ...) is saved as served. Existing files are overwritten. The call fails only if no file was saved.
| Name | Required | Description | Default |
|---|---|---|---|
| fit | No | Keep only the main content (drops menus, footers, and the like); falls back to the full page when nothing is left. Markdown only. | |
| urls | Yes | URLs to fetch (http or https), at most the server's per-call limit (see the server instructions); duplicates are fetched once. | |
| format | No | File format: 'markdown' (default), 'html', 'pdf' (page printed to PDF), 'screenshot' (PNG), 'mhtml' (single-file web archive), or 'raw' (the original response bytes, e.g. a PDF or image as served). | markdown |
| citations | No | Turn links into numbered references listed at the end. Markdown only. | |
| directory | No | Subdirectory of the server's download directory to save into; must stay inside it. Defaults to the download directory itself. | |
| timeout_s | No | Page load timeout per URL in seconds; defaults to the server setting. | |
| ignore_links | No | Drop links from the Markdown output. Markdown only. | |
| ignore_images | No | Drop images from the Markdown output. Markdown only. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Annotations carry no behavioral hints (readOnly=false, destructive=false), so description must carry the burden. It does: mentions overwrites, failure condition (fails only if no file saved), PDF transcription, and non-web handling. Good coverage, though it omits details like per-call URL limits (but schema covers that).
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
The description is a single, dense paragraph but front-loaded with the main action and result. It covers formats and key behaviors efficiently without fluff. Could be slightly more structured but is highly readable.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Given 8 parameters and no output schema, the description covers critical behaviors: success/failure, format specifics, and side effects. It does not explicitly state return format, but that's a minor gap. The 'raw' and PDF transcription details add completeness. Overall sufficient for an agent to call correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
Schema description coverage is 100%, so baseline is 3. The description adds context for 'raw' format and 'fit/citations' being Markdown-only, which is useful, but not extensive; the schema already explains each parameter's meaning.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
States a specific verb ('Fetch') and resource ('URLs') and explains it saves files to the server, returning paths. Explicitly distinguishes from sibling by listing supported formats and mention of server-side saving.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
Describes the tool's function and when it applies (fetching URLs to files), but does not explicitly contrast with the sibling 'fetch' tool. The 'raw' format and non-web content handling give context, but no direct 'use this instead of fetch' guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
fetchFetch web pagesARead-only
Fetch one or more web pages and return their content directly: Markdown (default), HTML, or a PNG screenshot. PDFs are transcribed to Markdown and images are returned as images. When several URLs are given, each result starts with a '' header. Failed URLs are reported as 'error: ...' lines; the call fails only if every URL fails. Use download to save other formats to files.
| Name | Required | Description | Default |
|---|---|---|---|
| fit | No | Keep only the main content (drops menus, footers, and the like); falls back to the full page when nothing is left. Markdown only. | |
| urls | Yes | URLs to fetch (http or https), at most the server's per-call limit (see the server instructions); duplicates are fetched once. | |
| format | No | Output format: 'markdown' (default), 'html' (rendered page HTML), or 'screenshot' (full-page PNG image). | markdown |
| citations | No | Turn links into numbered references listed at the end. Markdown only. | |
| timeout_s | No | Page load timeout per URL in seconds; defaults to the server setting. | |
| ignore_links | No | Drop links from the Markdown output. Markdown only. | |
| ignore_images | No | Drop images from the Markdown output. Markdown only. |
TDQS
Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?
Beyond the readOnly/openWorld annotations, the description discloses concrete non-obvious behavior: PDFs are transcribed to Markdown, images are returned as images, multi-URL results include header markers, failed URLs become 'error: ...' lines, and the call only fails if every URL fails. This is exactly the kind of behavioral detail an agent needs to interpret results.
Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.
Is the description appropriately sized, front-loaded, and free of redundancy?
Four sentences, each earning its place: core purpose and formats, special content handling, multi-URL/error semantics, and sibling routing. The description is front-loaded with the main action and contains no filler.
Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.
Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?
Even without an output schema, the description explains what the caller receives: direct content in the chosen format, result headers for multiple URLs, and error lines for failed URLs. Combined with the fully documented schema and read-only annotations, an agent has enough to invoke and interpret results correctly.
Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.
Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?
The input schema already describes every parameter in detail, so the baseline is 3. The description reinforces the format choices and adds context about PDF/image handling, but it does not add new parameter-level meaning beyond what the schema provides.
Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.
Does the description clearly state what the tool does and how it differs from similar tools?
The description opens with the verb 'Fetch' plus the resource 'web pages' and immediately enumerates the three output formats: Markdown, HTML, and PNG screenshot. It also distinguishes itself from the sibling tool by pointing to `download` for saving other formats, so an agent can tell them apart.
Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.
Does the description explain when to use this tool, when not to, or what alternatives exist?
It explicitly states when to use fetch for web content and routes other output needs to the `download` sibling. It also provides operational context for multi-URL calls, failure reporting, and the condition under which the call fails, giving an agent clear decision guidance.
Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.
Tool Schema Changelog
Recent tool additions, removals, and schema changes observed during successful MCP inspections.
2 tool updates
v0.3.0- First observed
download - First observed
fetch
TDQS
Scored across 2 tools
The two tools, `fetch` and `download`, have clearly distinct purposes: `fetch` returns content directly for immediate consumption, while `download` saves content to files. There is no ambiguity in their intended use cases, even though they share underlying capabilities.
Both tools use consistent short, verb-based names (`fetch` and `download`) that directly reflect their actions. They follow the same naming style (single-word imperative verbs), making them predictable and easy to remember.
With only two tools, the server feels very thin for a general-purpose web fetching service. While each tool is useful, the number is at the lower boundary, and many common operations (e.g., batch processing, URL validation, or content extraction options) are not represented as separate tools. This may be acceptable for a minimal server, but it limits agent flexibility.
The two tools cover basic fetching and downloading, but there are notable gaps: no way to list or manage previously downloaded files, no support for custom headers or POST requests, and no filtering or transformation of content beyond format selection. The surface is functional but misses common needs for a web-fetching domain.
Maintenance
Related MCP Connectors
Web data tools for AI agents: pages as markdown, search, maps, commerce, jobs, AI answers.
Fetch pages as markdown, search web and news, extract structured data. For AI agents.
Reliable web access for AI agents: smart HTTP, rotating proxies, and full-browser rendering.
Read a URL as clean markdown, screenshot a website, url to PDF. Web access for agents, no signup.
Related MCP Servers
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to clone entire websites, download files, manage authentication sessions, and analyze site information with support for JavaScript-heavy SPAs and dynamic content.9Apache 2.0
- AlicenseAqualityCmaintenanceEnables AI agents to read web pages reliably, returning clean markdown content, hyperlinks, and metadata without navigation or ad noise.36 npmMIT
- AlicenseAqualityDmaintenanceEnables AI agents to fetch any web page as clean markdown or screenshot it, turning URLs into LLM-ready context.211 npmMIT
- AlicenseNot gradedqualityCmaintenanceEnables AI agents to fetch and extract clean, readable content from web pages, and search within pages for specific queries, without needing a full browser.1MIT