Skip to main content
Glama

๐Ÿ”ฅ PyreCrawl โ€” Web Browsing Superpowers for Your AI Agent

License: MIT MCP Python 3.10+ PyPI

One command gives any AI agent the whole web. Scrape, extract, crawl, map, and search โ€” self-hosted, no API keys, no rate limits, no subscription.

PyreCrawl speaks MCP (Model Context Protocol), the standard tool interface for Claude, Cursor, VS Code, Codex, OpenCode, Hermes, and any MCP-compatible agent.

A smart auto-fallback ladder always picks the cheapest method that succeeds:

fast HTTP
    โ”‚  (403/503/Cloudflare challenge or empty body)
    โ–ผ
stealth browser (real Chromium + Cloudflare solver)
    โ”‚  (still blocked, or the page needs full JS rendering)
    โ–ผ
deep processing (LLM-ready markdown, citations, structured extraction)

โšก Tools exposed

Tool

What it does

scrape(url, prefer="auto")

Single URL โ†’ LLM-ready markdown

extract(url, schema)

Scrape + structured extraction (JsonCss schema)

map_site(root, include_pattern=None, limit=200)

Enumerate all internal URLs

crawl(root, max_pages=5, prefer="auto", include_paths=None, exclude_paths=None, max_depth=0)

Multi-page crawl with path filters + true BFS depth

document(url)

PDF/DOCX/PPTX โ†’ markdown (no browser, optional [docs] extras)

search(query, limit=10)

Web search via DuckDuckGo HTML (no API key)

search_papers(query, limit=8, source="arxiv", category=None)

Academic search via arXiv + Crossref (no API key) โ€” feed pdf_url into document

batch_scrape(urls[], ...)

Many URLs in ONE call โ€” parallel, deduped, cache-aware

deep_research(query, limit=5, scrape_top=3)

Search โ†’ evidence pack with [n] citations (no LLM synthesis โ€” your agent does that)

monitor(url, action, css_selector=None)

Change detection with persisted snapshots + unified diff

session(session, action, ...)

Persistent browser session (cookies kept) โ€” login walls, multi-step flows, screenshots

cache(action)

Inspect/clear/enable/disable the HTTP response cache

health()

Versions + import sanity check

MCP Resources (read-only state without a tool call): pyrecrawl://cache/stats ยท pyrecrawl://sessions ยท pyrecrawl://monitors

MCP Prompts (ready-made playbooks): research(topic) ยท rag_ingest(site) ยท watch_page(url)

Env flags

Variable

Default

Effect

PYRECRAWL_CACHE

off

1 = in-memory LRU (128 pages), or a directory path (reserved for disk mode)

PYRECRAWL_CACHE_TTL

900

Cache entry lifetime in seconds

PYRECRAWL_MONITOR_DIR

~/.pyrecrawl/monitors

Where monitor snapshots persist

prefer options: "auto" (default ladder) ยท "fast" (HTTP only) ยท "stealth" (CF bypass) ยท "llm" (deep processing).


๐Ÿš€ Install & Use (one-liner)

1. Install

UV (recommended โ€” one command, zero Python setup)

UV is a fast Python package manager that handles Python itself โ€” no need to install Python separately. Get it once:

# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh

# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"

Learn more about UV โ†’

Then run PyreCrawl directly โ€” no venv, no pip install, no Python download:

uvx pyrecrawl@latest
uv tool install pyrecrawl

Or via pipx (alternative)

pipx install pyrecrawl

Or via pip into a venv

pip install pyrecrawl

2. One-time browser engines

pyrecrawl setup

This installs Chromium + stealth browser engines (~2 min, one-time).

3. Register with your AI agent

# Auto-detect installed agents and write their MCP configs
pyrecrawl install

# Or target specific agents
pyrecrawl install claude-desktop cursor

# Dry-run to preview what would change
pyrecrawl install --dry-run

Supported agents: claude-desktop, claude-code, cursor, vscode, codex, opencode, hermes.

4. Start chatting

After installing + registering, restart your agent (or start a new session). Then ask:

"Scrape https://example.com and summarize it."

The tools appear as mcp_pyrecrawl_scrape, mcp_pyrecrawl_extract, mcp_pyrecrawl_map_site, mcp_pyrecrawl_crawl, mcp_pyrecrawl_search, mcp_pyrecrawl_health.


๐Ÿ“š Manual config (if pyrecrawl install doesn't match your setup)

Claude Desktop

Config file

  • Linux: ~/.config/Claude/claude_desktop_config.json

  • macOS: ~/Library/Application Support/Claude/claude_desktop_config.json

  • Windows: %AppData%\Claude\claude_desktop_config.json

{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

Claude Code

Config file: project-scoped .mcp.json

{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

Cursor

Config file: ~/.cursor/mcp.json

{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}

VS Code / Copilot

Config file: .vscode/mcp.json (project-scoped)

{
  "servers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"],
      "type": "stdio"
    }
  }
}

Codex CLI

Config file: ~/.codex/config.toml

[mcp_servers.pyrecrawl]
command = "uvx"
args = ["--from", "pyrecrawl", "pyrecrawl", "serve"]

OpenCode

Config file: ~/.config/opencode/opencode.json

{
  "mcp": {
    "pyrecrawl": {
      "type": "local",
      "command": ["uvx", "--from", "pyrecrawl", "pyrecrawl", "serve"],
      "enabled": true
    }
  }
}

Hermes

Config file

  • Linux/macOS: ~/.hermes/config.yaml

  • Windows: %LocalAppData%\hermes\config.yaml

mcp_servers:
  pyrecrawl:
    command: uvx
    args:
      - --from
      - pyrecrawl
      - pyrecrawl
      - serve
    enabled: true

Windows note: uvx must be on PATH. If not, use the full path to uvx.exe (e.g. C:\Users\<you>\AppData\Local\hermes\bin\uvx.exe).


๐Ÿง  How the ladder chooses

PyreCrawl runs each request through three tiers, stopping at the first one that returns a complete, LLM-ready result:

Concern

Fast tier

Stealth tier

Deep tier

Static HTML page

โœ… ~200ms

โ€”

โ€”

Cloudflare-protected

โŒ

โœ… Turnstile solver

โ€”

JS-heavy SPA

โŒ

โœ… real Chromium

โ€”

Live DOM data (input .value, JS state)

โŒ

โœ… js param

โ€”

LLM-ready markdown + citations

โ€”

โ€”

โœ… BM25, fit-markdown

Structured extraction (CSS schema)

โ€”

โ€”

โœ…

Deep crawl (BFS/DFS/BestFirst)

โ€”

โ€”

โœ… adaptive

The agent never has to pick. prefer="auto" does it every call.

Live DOM data with js and wait_for

Some sites keep the data you want in a DOM property (e.g. an <input>'s .value) that JS writes after an XHR โ€” it never appears in the serialized HTML. The scrape tool accepts two stealth-tier params for exactly this:

{
  "url": "https://temp-mail.org/id",
  "prefer": "stealth",
  "wait_for": "document.getElementById('mail').value.includes('@')",
  "js": "document.getElementById('mail').value"
}
  • wait_for โ€” a JS predicate expression polled until truthy (bounded by timeout). Use it instead of guessing a sleep for anything that arrives asynchronously.

  • js โ€” a JS expression evaluated once the page settles; the value comes back in meta.js_result. Errors are captured in meta.js_error (the page result is still returned, never a crash).


๐Ÿ“Š Compared to Firecrawl (hosted)

Firecrawl

PyreCrawl

Cost

Free 1k/mo, then $16โ€“333/mo

Free, self-hosted

Local LLM support

โŒ

โœ… Ollama / any LLM

Cloudflare bypass

โœ… (Fire-Engine, paid)

โœ… (free, built-in)

Markdown + BM25

โœ…

โœ…

Self-host

โŒ

โœ…

Academic paper search

โŒ

โœ… arXiv + Crossref (search_papers)

Hosted search API

โœ… /search

โš ๏ธ DuckDuckGo HTML + arXiv/Crossref (no key)


๐Ÿ”ง Development

git clone https://github.com/SanggonBoy/PyreCrawl.git
cd PyreCrawl
uv venv --python 3.12 .venv
source .venv/Scripts/activate  # Windows; or .venv/bin/activate on macOS/Linux
uv pip install -e ".[dev]"
python -m playwright install chromium
scrapling install

Run tests

python scripts/selfcheck.py   # real-network smoke test
python scripts/probe_stdio.py # stdio JSON-RPC probe

๐Ÿ“ฆ Publish

Maintainers only:

git tag v0.8.0
git push origin v0.8.0

GitHub Actions builds + uploads to PyPI via trusted publishing.


๐Ÿ”” Stay up to date

PyreCrawl checks PyPI on every startup and reports the latest version โ€” your MCP agent sees this automatically via the health() tool response and can notify you inline.

To check manually:

pyrecrawl version

To upgrade:

pyrecrawl update   # runs: uv tool upgrade pyrecrawl

Get notified of new releases: click Watch โ†’ Releases only at the GitHub repo to receive email notifications when a new version is published.


๐Ÿ“œ Uninstall

# Remove from all agent configs
pyrecrawl uninstall

# Remove the package
uv tool uninstall pyrecrawl

๐Ÿ›ก๏ธ License

MIT โ€” see LICENSE.

๐Ÿ™ Credits

Built on the shoulders of Scrapling and Crawl4AI โ€” both MIT, both excellent.