Skip to main content
Glama

Server Configuration

Describes the environment variables required to run the server.

NameRequiredDescriptionDefault
PYRECRAWL_CACHENo1 = in-memory LRU (128 pages), or a directory path (reserved for disk mode)off
PYRECRAWL_CACHE_TTLNoCache entry lifetime in seconds900
PYRECRAWL_MONITOR_DIRNoWhere monitor snapshots persist~/.pyrecrawl/monitors

Instructions

Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.

This server publishes no instructions, or was last inspected before Glama recorded them.

Capabilities

Features and capabilities supported by this server

Protocol revision2025-11-25

CapabilityDetails
tools
{
  "listChanged": false
}
prompts
{
  "listChanged": false
}
resources
{
  "subscribe": false,
  "listChanged": false
}
experimental
{}

Tools

Functions exposed to the LLM to take actions

NameDescription
scrapeA

Scrape a single URL → LLM-ready markdown.

Use this when the user shares a URL and wants its content (read, analyze, summarize, extract). Auto-escalates through fast→stealth→llm when blocked.

Args: url: Target URL (http/https). prefer: "auto" | "fast" | "stealth" | "llm". auto = fast first, escalate to stealth on block/short page. fast = cheap HTTP only (no JS). stealth = real Chromium + Cloudflare solver. llm = full Crawl4AI browser + BM25 fit-markdown. timeout: per-attempt timeout in seconds. include_html: include raw HTML in the response (large; off by default). js: (stealth only) JS expression evaluated against the live page after it settles. The value comes back in meta.js_result. Use for data that lives in DOM properties (e.g. an input's .value) rather than in serialized HTML. wait_for: (stealth only) JS predicate expression polled until truthy (bounded by timeout). Use to wait for content that arrives asynchronously after network_idle.

Returns: {url, final_url, status, markdown, title, method, elapsed_ms, meta} or {error, url, method} on failure.

extractA

Scrape + structured extraction using a CSS-based JSON schema.

Use when the user wants structured data (tables, lists, product info) extracted from a page. Define a CSS schema to target specific elements.

The schema is a JsonCssExtractionStrategy schema: { "name": "PageItems", "baseSelector": "div.item", "fields": [{"name": "title", "selector": "h2", "type": "text"}, ...] }

Returns parsed JSON in data.

map_siteA

Enumerate all internal URLs reachable from root.

Use when the user wants to map a site's structure or find all pages before crawling. Often paired with crawl or batch_scrape.

Args: root: Website root (e.g. "https://example.com/docs"). include_pattern: Optional regex; only URLs matching are returned. limit: Hard cap on returned URLs.

crawlA

Multi-page crawl: discover URLs on root, then scrape each.

Use this when the user wants to crawl an entire site section or docs, or needs multiple pages scraped in bulk. For single pages use scrape; for research questions use deep_research.

Args: root: start URL. max_pages: hard cap on pages scraped. css_selector: scope each page's html/markdown to the matched element (non-llm: lxml re-scope of the fetched HTML; llm: native crawl4ai css_selector). prefer: "auto" | "fast" | "stealth" | "llm" (llm = Crawl4AI BFS deep-crawl). include_paths: regex — keep only URLs matching (matched against full URL). exclude_paths: regex — drop URLs matching (e.g. /tag/|/page/\d+). max_depth: 0 = flat harvest from the root page's links (default); >0 = true BFS up to that link depth, honoring the filters.

Returns: {root, pages: [{url, markdown, title, ...}], count, discovered, elapsed_ms} or {error, root} on failure.

documentA

Extract text from a PDF/DOCX/PPTX URL → markdown (no browser).

Use when the user shares a link to a document (PDF, Word, PowerPoint) and wants its text content. Also useful after search_papers to get full text from a paper's pdf_url.

Content-type sniffed and routed to pypdf / python-docx / python-pptx. Optional deps — install with pip install 'pyrecrawl[docs]'.

searchA

Web search via DuckDuckGo HTML (no API key required).

Use this for targeted searches where you need anti-bot bypass (Cloudflare protection on DDG). For simple searches, the built-in web_search may suffice. For research questions, prefer deep_research (search + scrape + citations).

Returns [{url, title, snippet}, ...]. The smart ladder bypasses DDG's bot detection if needed.

Returns: {query, results: [{url, title, snippet}], count} or {error, query} on failure.

batch_scrapeA

Scrape MANY URLs in ONE call (parallel, deduped, cache-aware).

Use when the user provides multiple URLs or you have a list of pages to fetch. More efficient than calling scrape N times.

Args: urls: Target URLs (deduped automatically; empties dropped). prefer: "auto" | "fast" | "stealth" | "llm". timeout: per-URL timeout in seconds. max_concurrency: parallel workers (default 4). include_html: include raw HTML per result (large; off by default).

Returns {requested, unique, succeeded, failed, results[]}. Per-URL failures are isolated — other URLs still succeed.

Returns: {requested, unique, succeeded, failed, results: [{url, markdown, ...}]}

deep_researchA

Search the web, then pull the top sources as EVIDENCE (no LLM synthesis).

PRIMARY RESEARCH TOOL — use when the user asks to research, investigate, deep-dive, fact-check, or learn about a topic. Returns a citations list with stable [n] numbers and an evidence list of per-source markdown — the agent does the synthesis from evidence.

Multi-pass mode: set iterations=2-3 to auto-run additional searches with refined queries (alternatives, criticism, latest developments) and append deduplicated evidence. Each pass adds up to scrape_top new sources.

Args: query: search string. limit: how many search results to fetch. scrape_top: how many of those to actually fetch content from. prefer: "auto" | "fast" | "stealth" | "llm". iterations: 1 (default, single pass), 2-3 (multi-pass with refined queries targeting evidence gaps). Each pass searches from a different angle and deduplicates by URL.

Returns: {query, iterations_run, queries: [str], hits: [{url, title, snippet}], citations: [{url, title}], evidence: [{url, title, markdown}], scraped, used_engines, elapsed_ms} or {error, query, hint} on failure.

monitorA

Track a URL over time and report meaningful content changes.

Use when the user wants to watch a page for updates (price changes, new blog posts, status updates). Ask me to 'set up monitoring for ' for a guided setup playbook.

Args: url: target URL. action: "check" | "history" | "forget". prefer: ladder preference, same as scrape. css_selector: scope the diff to one element (so banner / nav changes don't trigger false positives).

Snapshots persist under PYRECRAWL_MONITOR_DIR (default ~/.pyrecrawl/monitors/). check returns status of new | unchanged | changed | error and a unified diff when the page changed.

Returns: {url, status: "new"|"unchanged"|"changed"|"error", diff?: str, snapshot_chars?: int, elapsed_ms?: int}

search_papersA

Search academic papers via arXiv or Crossref — no API keys.

Args: query: free-text search, e.g. "transformer attention scaling laws". limit: max results (1-25 arXiv / 1-20 crossref). source: "arxiv" (CS/physics/math preprints, default) or "crossref" (all fields, DOI-backed). category: optional arXiv category filter, e.g. "cs.LG", "cs.CV".

Returns papers with id/url/pdf_url/title/authors/summary/published. Feed pdf_url into the document tool to extract full text.

Returns: {query, source, papers: [{title, authors, abstract, url, pdf_url?, ...}]} or {error, query} on failure.

cacheA

Inspect the response cache: stats, clear, enable, or disable.

Use this to check cache hit rates before large batch jobs, or to clear stale cached responses when a site's content has changed.

Args: action: "stats" (default) | "clear" | "disable" | "enable".

healthA

Verify PyreCrawl is working: engine versions, dependencies, update status.

Use this at the start of a session or before a large scraping job to confirm all engines are installed and up to date.

sessionA

Drive a persistent browser session — cookies & JS state kept across calls.

Use for login walls and multi-step flows the one-shot ladder can't handle (e.g. user needs to log in first, then scrape a protected page). For simple pages use scrape; for research use deep_research.

One-shot scrape has no session memory; here each action runs against the same live page.

Args: session: named session; reuse the same name to keep state. action: "open" (url) — navigate, returns url/title/status "click" (selector) — click an element "fill" (selector, text) — type into an input "type" (key) — press a key, e.g. "Enter" "eval" (js) — run a JS expression, returns value "wait" (selector?) — wait for selector or sleep timeout_ms "content" — url/title/visible text of current page "screenshot" (full_page?) — returns png_base64 "cookies" — list session cookies "close" — destroy the session "list" — show live sessions Returns {session, action, ...result, elapsed_ms}; errors as {error}.

Prompts

Interactive templates invoked by user choice

NameDescription
researchPlaybook: evidence-first research on any topic using PyreCrawl tools. Trigger phrases: 'research X', 'investigate X', 'deep dive into X', 'tell me about X with sources', 'fact check X'.
rag_ingestPlaybook: turn a site into clean markdown for a RAG index. Trigger phrases: 'prepare site for RAG', 'ingest site into index', 'crawl site for knowledge base'.
watch_pagePlaybook: set up change watching on a page with a sensible baseline. Trigger phrases: 'monitor this page', 'watch for changes on URL', 'notify me when this updates', 'track this URL'.

Resources

Contextual data attached and managed by the client

NameDescription
resource_cache_statsResponse cache state: enabled, items, hits/misses, TTL.
resource_sessionsLive browser sessions with idle time.
resource_monitorsMonitored URLs and their last check (new/unchanged/changed state).

TDQS

A4.4/5.0

Scored across 13 tools

Disambiguation5/5

Each tool has a clearly distinct trigger: single-page scrape, structured extraction, bulk scrape, site crawl, document parsing, web search, deep research, academic search, monitoring, sessions, and cache/health maintenance. Even though several tools fetch pages, their arguments and return shapes make the intended use obvious.

Naming Consistency4/5

Names are all lowercase and action-oriented, but the convention is mixed: single-word verbs (scrape, crawl, search), verb_noun compounds (map_site, search_papers), and modifier compounds (batch_scrape, deep_research), plus noun-style tools like health and session. This is readable and mostly predictable, but not a uniform pattern.

Tool Count5/5

13 tools is well within the ideal range for a scraping and research suite. Each tool covers a distinct operation, and the utility tools (cache, health) support the workflow without feeling like filler.

Completeness5/5

The surface covers the full scraping/research workflow: single and batch scraping, crawling, structured extraction, document parsing, search, deep research, academic papers, monitoring, sessions, and cache/health management. There are no obvious dead ends — monitors can be forgotten, sessions can be closed, and cache can be cleared or disabled.

Maintenance

ActivityMaintained
ResponsivenessNo issues