PyreCrawl
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| PYRECRAWL_CACHE | No | 1 = in-memory LRU (128 pages), or a directory path (reserved for disk mode) | off |
| PYRECRAWL_CACHE_TTL | No | Cache entry lifetime in seconds | 900 |
| PYRECRAWL_MONITOR_DIR | No | Where monitor snapshots persist | ~/.pyrecrawl/monitors |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| scrapeA | Scrape a single URL → LLM-ready markdown. Use this when the user shares a URL and wants its content (read, analyze, summarize, extract). Auto-escalates through fast→stealth→llm when blocked. Args:
url: Target URL (http/https).
prefer: "auto" | "fast" | "stealth" | "llm".
auto = fast first, escalate to stealth on block/short page.
fast = cheap HTTP only (no JS).
stealth = real Chromium + Cloudflare solver.
llm = full Crawl4AI browser + BM25 fit-markdown.
timeout: per-attempt timeout in seconds.
include_html: include raw HTML in the response (large; off by default).
js: (stealth only) JS expression evaluated against the live page
after it settles. The value comes back in Returns: {url, final_url, status, markdown, title, method, elapsed_ms, meta} or {error, url, method} on failure. |
| extractA | Scrape + structured extraction using a CSS-based JSON schema. Use when the user wants structured data (tables, lists, product info) extracted from a page. Define a CSS schema to target specific elements. The schema is a JsonCssExtractionStrategy schema: { "name": "PageItems", "baseSelector": "div.item", "fields": [{"name": "title", "selector": "h2", "type": "text"}, ...] } Returns parsed JSON in |
| map_siteA | Enumerate all internal URLs reachable from Use when the user wants to map a site's structure or find all pages
before crawling. Often paired with Args: root: Website root (e.g. "https://example.com/docs"). include_pattern: Optional regex; only URLs matching are returned. limit: Hard cap on returned URLs. |
| crawlA | Multi-page crawl: discover URLs on Use this when the user wants to crawl an entire site section or docs,
or needs multiple pages scraped in bulk. For single pages use Args:
root: start URL.
max_pages: hard cap on pages scraped.
css_selector: scope each page's html/markdown to the matched element
(non-llm: lxml re-scope of the fetched HTML; llm: native crawl4ai
css_selector).
prefer: "auto" | "fast" | "stealth" | "llm" (llm = Crawl4AI BFS deep-crawl).
include_paths: regex — keep only URLs matching (matched against full URL).
exclude_paths: regex — drop URLs matching (e.g. Returns: {root, pages: [{url, markdown, title, ...}], count, discovered, elapsed_ms} or {error, root} on failure. |
| documentA | Extract text from a PDF/DOCX/PPTX URL → markdown (no browser). Use when the user shares a link to a document (PDF, Word, PowerPoint)
and wants its text content. Also useful after Content-type sniffed and routed to pypdf / python-docx / python-pptx.
Optional deps — install with |
| searchA | Web search via DuckDuckGo HTML (no API key required). Use this for targeted searches where you need anti-bot bypass (Cloudflare
protection on DDG). For simple searches, the built-in web_search may suffice.
For research questions, prefer Returns [{url, title, snippet}, ...]. The smart ladder bypasses DDG's bot detection if needed. Returns: {query, results: [{url, title, snippet}], count} or {error, query} on failure. |
| batch_scrapeA | Scrape MANY URLs in ONE call (parallel, deduped, cache-aware). Use when the user provides multiple URLs or you have a list of pages
to fetch. More efficient than calling Args: urls: Target URLs (deduped automatically; empties dropped). prefer: "auto" | "fast" | "stealth" | "llm". timeout: per-URL timeout in seconds. max_concurrency: parallel workers (default 4). include_html: include raw HTML per result (large; off by default). Returns {requested, unique, succeeded, failed, results[]}. Per-URL failures are isolated — other URLs still succeed. Returns: {requested, unique, succeeded, failed, results: [{url, markdown, ...}]} |
| deep_researchA | Search the web, then pull the top sources as EVIDENCE (no LLM synthesis). PRIMARY RESEARCH TOOL — use when the user asks to research, investigate,
deep-dive, fact-check, or learn about a topic. Returns a Multi-pass mode: set iterations=2-3 to auto-run additional searches with
refined queries (alternatives, criticism, latest developments) and append
deduplicated evidence. Each pass adds up to Args: query: search string. limit: how many search results to fetch. scrape_top: how many of those to actually fetch content from. prefer: "auto" | "fast" | "stealth" | "llm". iterations: 1 (default, single pass), 2-3 (multi-pass with refined queries targeting evidence gaps). Each pass searches from a different angle and deduplicates by URL. Returns: {query, iterations_run, queries: [str], hits: [{url, title, snippet}], citations: [{url, title}], evidence: [{url, title, markdown}], scraped, used_engines, elapsed_ms} or {error, query, hint} on failure. |
| monitorA | Track a URL over time and report meaningful content changes. Use when the user wants to watch a page for updates (price changes, new blog posts, status updates). Ask me to 'set up monitoring for ' for a guided setup playbook. Args:
url: target URL.
action: "check" | "history" | "forget".
prefer: ladder preference, same as Snapshots persist under Returns: {url, status: "new"|"unchanged"|"changed"|"error", diff?: str, snapshot_chars?: int, elapsed_ms?: int} |
| search_papersA | Search academic papers via arXiv or Crossref — no API keys. Args: query: free-text search, e.g. "transformer attention scaling laws". limit: max results (1-25 arXiv / 1-20 crossref). source: "arxiv" (CS/physics/math preprints, default) or "crossref" (all fields, DOI-backed). category: optional arXiv category filter, e.g. "cs.LG", "cs.CV". Returns papers with id/url/pdf_url/title/authors/summary/published.
Feed pdf_url into the Returns: {query, source, papers: [{title, authors, abstract, url, pdf_url?, ...}]} or {error, query} on failure. |
| cacheA | Inspect the response cache: stats, clear, enable, or disable. Use this to check cache hit rates before large batch jobs, or to clear stale cached responses when a site's content has changed. Args: action: "stats" (default) | "clear" | "disable" | "enable". |
| healthA | Verify PyreCrawl is working: engine versions, dependencies, update status. Use this at the start of a session or before a large scraping job to confirm all engines are installed and up to date. |
| sessionA | Drive a persistent browser session — cookies & JS state kept across calls. Use for login walls and multi-step flows the one-shot ladder can't handle
(e.g. user needs to log in first, then scrape a protected page).
For simple pages use One-shot scrape has no session memory; here each action runs against the same live page. Args: session: named session; reuse the same name to keep state. action: "open" (url) — navigate, returns url/title/status "click" (selector) — click an element "fill" (selector, text) — type into an input "type" (key) — press a key, e.g. "Enter" "eval" (js) — run a JS expression, returns value "wait" (selector?) — wait for selector or sleep timeout_ms "content" — url/title/visible text of current page "screenshot" (full_page?) — returns png_base64 "cookies" — list session cookies "close" — destroy the session "list" — show live sessions Returns {session, action, ...result, elapsed_ms}; errors as {error}. |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
| research | Playbook: evidence-first research on any topic using PyreCrawl tools. Trigger phrases: 'research X', 'investigate X', 'deep dive into X', 'tell me about X with sources', 'fact check X'. |
| rag_ingest | Playbook: turn a site into clean markdown for a RAG index. Trigger phrases: 'prepare site for RAG', 'ingest site into index', 'crawl site for knowledge base'. |
| watch_page | Playbook: set up change watching on a page with a sensible baseline. Trigger phrases: 'monitor this page', 'watch for changes on URL', 'notify me when this updates', 'track this URL'. |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
| resource_cache_stats | Response cache state: enabled, items, hits/misses, TTL. |
| resource_sessions | Live browser sessions with idle time. |
| resource_monitors | Monitored URLs and their last check (new/unchanged/changed state). |
TDQS
Scored across 13 tools
Each tool has a clearly distinct trigger: single-page scrape, structured extraction, bulk scrape, site crawl, document parsing, web search, deep research, academic search, monitoring, sessions, and cache/health maintenance. Even though several tools fetch pages, their arguments and return shapes make the intended use obvious.
Names are all lowercase and action-oriented, but the convention is mixed: single-word verbs (scrape, crawl, search), verb_noun compounds (map_site, search_papers), and modifier compounds (batch_scrape, deep_research), plus noun-style tools like health and session. This is readable and mostly predictable, but not a uniform pattern.
13 tools is well within the ideal range for a scraping and research suite. Each tool covers a distinct operation, and the utility tools (cache, health) support the workflow without feeling like filler.
The surface covers the full scraping/research workflow: single and batch scraping, crawling, structured extraction, document parsing, search, deep research, academic papers, monitoring, sessions, and cache/health management. There are no obvious dead ends — monitors can be forgotten, sessions can be closed, and cache can be cleared or disabled.