Skip to main content
Glama
README.md
# ๐Ÿ”ฅ PyreCrawl โ€” Web Browsing Superpowers for Your AI Agent

[![License: MIT](https://img.shields.io/badge/License-MIT-yellow.svg)](https://opensource.org/licenses/MIT)
[![MCP](https://img.shields.io/badge/MCP-1.0-blue.svg)](https://modelcontextprotocol.io/)
[![Python 3.10+](https://img.shields.io/badge/python-3.10+-blue.svg)](https://www.python.org/)
[![PyPI](https://img.shields.io/pypi/v/pyrecrawl.svg)](https://pypi.org/project/pyrecrawl/)
[![GitHub stars](https://img.shields.io/github/stars/SanggonBoy/PyreCrawl?logo=github)](https://github.com/SanggonBoy/PyreCrawl/stargazers)
[![Downloads / 30d](https://img.shields.io/endpoint?url=https%3A%2F%2Fpyrecrawl-stats.fajarnugraha90543.workers.dev%2Fbadge%2Fdownloads)](https://pypistats.org/packages/pyrecrawl)
[![Active users / 30d](https://img.shields.io/endpoint?url=https%3A%2F%2Fpyrecrawl-stats.fajarnugraha90543.workers.dev%2Fbadge%2Fusers)](https://github.com/SanggonBoy/PyreCrawl#-privacy--anonymous-usage-ping)

**One command gives any AI agent the whole web.** Scrape, extract, crawl, map, and search โ€”
self-hosted, no API keys, no rate limits, no subscription.

PyreCrawl speaks **MCP** (Model Context Protocol), the standard tool interface for Claude,
Cursor, VS Code, Codex, OpenCode, Hermes, and any MCP-compatible agent.

A **smart auto-fallback ladder** always picks the cheapest method that succeeds:

```
fast HTTP
    โ”‚  (403/503/Cloudflare challenge or empty body)
    โ–ผ
stealth browser (real Chromium + Cloudflare solver)
    โ”‚  (still blocked, or the page needs full JS rendering)
    โ–ผ
deep processing (LLM-ready markdown, citations, structured extraction)
```

## โšก Tools exposed

| Tool | What it does |
|---|---|
| `scrape(url, prefer="auto")` | Single URL โ†’ LLM-ready markdown |
| `extract(url, schema)` | Scrape + structured extraction (JsonCss schema) |
| `map_site(root, include_pattern=None, limit=200)` | Enumerate all internal URLs |
| `crawl(root, max_pages=5, prefer="auto", include_paths=None, exclude_paths=None, max_depth=0)` | Multi-page crawl with path filters + true BFS depth |
| `document(url)` | PDF/DOCX/PPTX โ†’ markdown (no browser, optional `[docs]` extras) |
| `search(query, limit=10)` | Web search via DuckDuckGo HTML (no API key) |
| `search_papers(query, limit=8, source="arxiv", category=None)` | Academic search via arXiv + Crossref (no API key) โ€” feed `pdf_url` into `document` |
| `batch_scrape(urls[], ...)` | Many URLs in ONE call โ€” parallel, deduped, cache-aware |
| `deep_research(query, limit=5, scrape_top=3)` | Search โ†’ evidence pack with [n] citations (no LLM synthesis โ€” your agent does that) |
| `monitor(url, action, css_selector=None)` | Change detection with persisted snapshots + unified diff |
| `session(session, action, ...)` | Persistent browser session (cookies kept) โ€” login walls, multi-step flows, screenshots |
| `cache(action)` | Inspect/clear/enable/disable the HTTP response cache |
| `health()` | Versions + import sanity check |

**MCP Resources** (read-only state without a tool call):
`pyrecrawl://cache/stats` ยท `pyrecrawl://sessions` ยท `pyrecrawl://monitors`

**MCP Prompts** (ready-made playbooks): `research(topic)` ยท `rag_ingest(site)` ยท `watch_page(url)`

### Env flags

| Variable | Default | Effect |
|---|---|---|
| `PYRECRAWL_CACHE` | off | `1` = in-memory LRU (128 pages), or a directory path (reserved for disk mode) |
| `PYRECRAWL_CACHE_TTL` | `900` | Cache entry lifetime in seconds |
| `PYRECRAWL_MONITOR_DIR` | `~/.pyrecrawl/monitors` | Where monitor snapshots persist |
| `PYRECRAWL_NO_TELEMETRY` | off | `1` = disable the anonymous startup ping (also honors `DO_NOT_TRACK=1`) |

`prefer` options: `"auto"` (default ladder) ยท `"fast"` (HTTP only) ยท `"stealth"` (CF bypass) ยท `"llm"` (deep processing).

---

## ๐Ÿš€ Install & Use (one-liner)

### 1. Install

#### [UV](https://docs.astral.sh/uv/) (recommended โ€” one command, zero Python setup)

UV is a fast Python package manager that handles Python itself โ€”
no need to install Python separately. Get it once:

```bash
# macOS / Linux
curl -LsSf https://astral.sh/uv/install.sh | sh

# Windows (PowerShell)
powershell -ExecutionPolicy ByPass -c "irm https://astral.sh/uv/install.ps1 | iex"
```

[Learn more about UV โ†’](https://docs.astral.sh/uv/)

Then run PyreCrawl directly โ€” no venv, no `pip install`, no Python download:

```bash
uvx pyrecrawl@latest
```

#### Or via uv tool install (persistent, recommended for regular use)

```bash
uv tool install pyrecrawl
```

#### Or via pipx (alternative)

```bash
pipx install pyrecrawl
```

#### Or via pip into a venv

```bash
pip install pyrecrawl
```

### 2. One-time browser engines

```bash
pyrecrawl setup
```

This installs Chromium + stealth browser engines (~2 min, one-time).

### 3. Register with your AI agent

```bash
# Auto-detect installed agents and write their MCP configs
pyrecrawl install

# Or target specific agents
pyrecrawl install claude-desktop cursor

# Dry-run to preview what would change
pyrecrawl install --dry-run
```

Supported agents: `claude-desktop`, `claude-code`, `cursor`, `vscode`, `codex`, `opencode`, `hermes`.

### 4. Start chatting

After installing + registering, **restart your agent** (or start a new session). Then ask:

> *"Scrape https://example.com and summarize it."*

Available tools:

| Tool | What it does |
|------|-------------|
| `scrape` | Fetch a single URL โ†’ markdown (auto-escalates past Cloudflare) |
| `extract` | Scrape + structured extraction via CSS schema โ†’ JSON |
| `map_site` | Enumerate all internal URLs from a root |
| `crawl` | Multi-page crawl: discover + scrape in bulk |
| `batch_scrape` | Fetch many URLs in one parallel call |
| `search` | Web search via DuckDuckGo with anti-bot bypass |
| `search_papers` | Academic paper search (arXiv / Crossref) |
| `deep_research` | Search + scrape + citations in one call โ€” **primary research tool** |
| `document` | Extract text from PDF/DOCX/PPTX URLs |
| `monitor` | Track a URL for content changes over time |
| `session` | Persistent browser session for login walls |
| `cache` | Inspect or clear the response cache |
| `health` | Verify engine availability + version |

Plus 3 guided prompts: `research`, `rag_ingest`, `watch_page`.

### Quick examples

Ask your agent naturally โ€” no special syntax needed:

| You say | Agent uses |
|---------|-----------|
| *"Scrape https://example.com and summarize it"* | `scrape` โ†’ returns markdown โ†’ agent summarizes |
| *"Research Rust memory safety vulnerabilities"* | `deep_research` โ†’ search + scrape + citations |
| *"Deep research on AI regulation worldwide"* | `deep_research(iterations=3)` โ†’ multi-pass with refined queries |
| *"Extract all product names and prices from this page"* | `extract` โ†’ CSS schema โ†’ structured JSON |
| *"Crawl https://docs.example.com and give me an overview"* | `crawl` โ†’ multi-page โ†’ summary |
| *"Monitor this page for price changes"* | `monitor` โ†’ baseline snapshot โ†’ periodic diff |
| *"Find papers about transformer attention"* | `search_papers` โ†’ arXiv results |
| *"What's the current cache hit rate?"* | `cache` โ†’ stats |

---

## ๐Ÿง  Skills โ€” Maximize Your Agent's Research Quality

PyreCrawl tools give your agent **hands** (scrape, crawl, search). But the agent still needs a **brain** โ€” instructions on *when* to use which tool, *how* to chain research passes, and *what* anti-hallucination rules to follow.

That's what **[PyreCrawl Skills](https://github.com/SanggonBoy/pyrecrawl-skills)** provides.

| | MCP Tools (this repo) | Skills ([pyrecrawl-skills](https://github.com/SanggonBoy/pyrecrawl-skills)) |
|---|---|---|
| **Role** | Execute web operations | Tell the agent how to use them |
| **Analogy** | Hands | Brain |
| **Example** | `deep_research(query, iterations=3)` | "Run 3 passes, check gaps after each, cite everything" |
| **Required?** | Yes (the engine) | Optional (but recommended for research quality) |

**Quick setup:**
```bash
# 1. Install the tools (you already have this)
uvx pyrecrawl@latest

# 2. Add the research skill to your project
git clone https://github.com/SanggonBoy/pyrecrawl-skills.git /tmp/pyrecrawl-skills
cp /tmp/pyrecrawl-skills/pyrecrawl-research/SKILL.md ./CLAUDE.md  # or .cursorrules / AGENTS.md
```

> **Without skills:** Your agent has powerful tools but improvises usage.
> **With skills:** Your agent follows a proven research protocol with anti-hallucination guardrails.

---

## ๐Ÿ“š Manual config (if `pyrecrawl install` doesn't match your setup)

### Claude Desktop

**Config file**
- Linux: `~/.config/Claude/claude_desktop_config.json`
- macOS: `~/Library/Application Support/Claude/claude_desktop_config.json`
- Windows: `%AppData%\Claude\claude_desktop_config.json`

```json
{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}
```

### Claude Code

**Config file**: project-scoped `.mcp.json`

```json
{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}
```

### Cursor

**Config file**: `~/.cursor/mcp.json`

```json
{
  "mcpServers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"]
    }
  }
}
```

### VS Code / Copilot

**Config file**: `.vscode/mcp.json` (project-scoped)

```json
{
  "servers": {
    "pyrecrawl": {
      "command": "uvx",
      "args": ["--from", "pyrecrawl", "pyrecrawl", "serve"],
      "type": "stdio"
    }
  }
}
```

### Codex CLI

**Config file**: `~/.codex/config.toml`

```toml
[mcp_servers.pyrecrawl]
command = "uvx"
args = ["--from", "pyrecrawl", "pyrecrawl", "serve"]
```

### OpenCode

**Config file**: `~/.config/opencode/opencode.json`

```json
{
  "mcp": {
    "pyrecrawl": {
      "type": "local",
      "command": ["uvx", "--from", "pyrecrawl", "pyrecrawl", "serve"],
      "enabled": true
    }
  }
}
```

### Hermes

**Config file**
- Linux/macOS: `~/.hermes/config.yaml`
- Windows: `%LocalAppData%\hermes\config.yaml`

```yaml
mcp_servers:
  pyrecrawl:
    command: uvx
    args:
      - --from
      - pyrecrawl
      - pyrecrawl
      - serve
    enabled: true
```

> **Windows note:** `uvx` must be on PATH. If not, use the full path to `uvx.exe` (e.g. `C:\Users\<you>\AppData\Local\hermes\bin\uvx.exe`).

---

## ๐Ÿง  How the ladder chooses

PyreCrawl runs each request through three tiers, stopping at the first one that returns
a complete, LLM-ready result:

| Concern | Fast tier | Stealth tier | Deep tier |
|---|---|---|---|
| Static HTML page | โœ… ~200ms | โ€” | โ€” |
| Cloudflare-protected | โŒ | โœ… Turnstile solver | โ€” |
| JS-heavy SPA | โŒ | โœ… real Chromium | โ€” |
| Live DOM data (input `.value`, JS state) | โŒ | โœ… `js` param | โ€” |
| LLM-ready markdown + citations | โ€” | โ€” | โœ… BM25, fit-markdown |
| Structured extraction (CSS schema) | โ€” | โ€” | โœ… |
| Deep crawl (BFS/DFS/BestFirst) | โ€” | โ€” | โœ… adaptive |

The agent never has to pick. `prefer="auto"` does it every call.

### Live DOM data with `js` and `wait_for`

Some sites keep the data you want in a DOM *property* (e.g. an `<input>`'s `.value`)
that JS writes after an XHR โ€” it never appears in the serialized HTML. The
`scrape` tool accepts two stealth-tier params for exactly this:

```json
{
  "url": "https://temp-mail.org/id",
  "prefer": "stealth",
  "wait_for": "document.getElementById('mail').value.includes('@')",
  "js": "document.getElementById('mail').value"
}
```

- `wait_for` โ€” a JS **predicate expression** polled until truthy (bounded by `timeout`).
  Use it instead of guessing a sleep for anything that arrives asynchronously.
- `js` โ€” a JS **expression** evaluated once the page settles; the value comes back
  in `meta.js_result`. Errors are captured in `meta.js_error` (the page result is
  still returned, never a crash).

---

## ๐Ÿ“Š Compared to Firecrawl (hosted)

| | Firecrawl | PyreCrawl |
|---|---|---|
| Cost | Free 1k/mo, then $16โ€“333/mo | **Free, self-hosted** |
| Local LLM support | โŒ | โœ… Ollama / any LLM |
| Cloudflare bypass | โœ… (Fire-Engine, paid) | โœ… (free, built-in) |
| Markdown + BM25 | โœ… | โœ… |
| Self-host | โŒ | โœ… |
| Academic paper search | โŒ | โœ… arXiv + Crossref (`search_papers`) |
| Hosted search API | โœ… /search | โš ๏ธ DuckDuckGo HTML + arXiv/Crossref (no key) |

---

## ๐Ÿ”ง Development

```bash
git clone https://github.com/SanggonBoy/PyreCrawl.git
cd PyreCrawl
uv venv --python 3.12 .venv
source .venv/Scripts/activate  # Windows; or .venv/bin/activate on macOS/Linux
uv pip install -e ".[dev]"
python -m playwright install chromium
scrapling install
```

### Run tests

```bash
python scripts/selfcheck.py        # real-network smoke test (13 tools + engines)
python scripts/probe_stdio.py      # stdio JSON-RPC probe
python scripts/test_ladder_bug.py  # SPA-shell ladder escalation regression
python scripts/test_js_eval.py     # stealth js/wait_for params regression
python scripts/test_scope_selector.py  # crawl css_selector/max_depth wiring
python scripts/test_link_harvest.py    # map/BFS link purity regression
```

---

## ๐Ÿ“ฆ Publish

Maintainers only:

```bash
git tag vX.Y.Z
git push origin vX.Y.Z
```

GitHub Actions builds + uploads to PyPI via [trusted publishing](https://docs.pypi.org/trusted-publishers/).

---

## ๐Ÿ”” Stay up to date

PyreCrawl checks PyPI on every startup and reports the latest version โ€” your
MCP agent sees this automatically via the `health()` tool response and can
notify you inline.

To check manually:

```bash
pyrecrawl version
```

To upgrade:

```bash
pyrecrawl update   # runs: uv tool upgrade pyrecrawl
```

**Get notified of new releases:** click **Watch** โ†’ **Releases only** at the
[GitHub repo](https://github.com/SanggonBoy/PyreCrawl) to receive email
notifications when a new version is published.

---

> [!NOTE]
> PyreCrawl sends **one anonymous usage ping per 24 h** at server startup โ€” see
> [Privacy](#-privacy--anonymous-usage-ping) for exactly what's sent and how to opt out.

## ๐Ÿ”’ Privacy โ€” anonymous usage ping

PyreCrawl phones home **once per 24 h** with a tiny anonymous ping when the MCP
server starts, so we can count real users (DAU/MAU) instead of raw downloads.

| Sent (4 fields, ~100 bytes) | Never sent |
|---|---|
| Hashed machine id (SHA-256 of hostname+MAC โ€” not reversible) | Your IP (not stored) |
| PyreCrawl version | Any URL you scrape |
| Python version | Any page content or search queries |
| OS family (`windows` / `linux` / `darwin`) | Anything else |

Client code: [`src/pyrecrawl/telemetry.py`](src/pyrecrawl/telemetry.py) (~90 lines, stdlib only) ยท
Collector: [`workers/telemetry/`](workers/telemetry/) โ€” a self-hostable Cloudflare Worker + D1, no third-party analytics service.

Opt out any time:

```bash
export PYRECRAWL_NO_TELEMETRY=1   # or the industry-standard DO_NOT_TRACK=1
```

---

## ๐Ÿ“œ Uninstall

```bash
# Remove from all agent configs
pyrecrawl uninstall

# Remove the package
uv tool uninstall pyrecrawl
```

---

## ๐Ÿ›ก๏ธ License

MIT โ€” see [LICENSE](LICENSE).

<!-- mcp-name: io.github.SanggonBoy/PyreCrawl -->

TDQS

A4.4/5.0

Scored across 13 tools

Disambiguation5/5

Each tool has a clearly distinct trigger: single-page scrape, structured extraction, bulk scrape, site crawl, document parsing, web search, deep research, academic search, monitoring, sessions, and cache/health maintenance. Even though several tools fetch pages, their arguments and return shapes make the intended use obvious.

Naming Consistency4/5

Names are all lowercase and action-oriented, but the convention is mixed: single-word verbs (scrape, crawl, search), verb_noun compounds (map_site, search_papers), and modifier compounds (batch_scrape, deep_research), plus noun-style tools like health and session. This is readable and mostly predictable, but not a uniform pattern.

Tool Count5/5

13 tools is well within the ideal range for a scraping and research suite. Each tool covers a distinct operation, and the utility tools (cache, health) support the workflow without feeling like filler.

Completeness5/5

The surface covers the full scraping/research workflow: single and batch scraping, crawling, structured extraction, document parsing, search, deep research, academic papers, monitoring, sessions, and cache/health management. There are no obvious dead ends โ€” monitors can be forgotten, sessions can be closed, and cache can be cleared or disabled.

Maintenance

ActivityMaintained
ResponsivenessNo issues