Skip to main content
Glama

searxNcrawl

MCP server and CLI toolkit for web search and crawling, built on Crawl4AI and SearXNG.

Published at github.com/DasDigitaleMomentum/searxNcrawl — maintained by DDM – Das Digitale Momentum GmbH & Co KG. Successor to searxng-mcp.

Quick Start

Pick your setup:

Docker Compose

MCP server with Playwright/Chromium, ready in one command. SearXNG required separately for search.

cp .env.example .env          # set SEARXNG_URL to your SearXNG instance
docker compose up --build

➜ MCP server at http://localhost:9555/mcp

pip (standalone)

CLI tools, Python API, and MCP server. SearXNG required for search.

python -m venv .venv && source .venv/bin/activate
pip install -e .
playwright install chromium

uv (standalone)

Same capabilities as pip.

uv sync
uv run playwright install chromium

What you get

Feature

Docker Compose

pip / uv

MCP Server (STDIO)

MCP Server (HTTP)

Web Crawl

Web Search

✅¹

✅¹

CLI Tools

via exec²

Python API

CORS (HTTP)

¹ Requires a SearXNG instance. ² docker compose exec searxncrawl crawl ...

Related MCP server: evo-scry

Features

Crawling

  • Single page, multi-page, and site crawling (DFS with depth/page limits)

  • Production-tested extraction config optimized for documentation sites

  • Configurable timeouts with graceful error handling

Content Quality

  • Markdown deduplicationexact (default) removes repeated blocks, off disables it

  • Link removal — strip all links for cleaner LLM context (--remove-links)

  • Dedup guardrails — non-destructive metadata signals when removal is unusually aggressive

  • SearXNG metasearch integration (privacy-respecting)

  • Configurable language, time range, categories, engines, safe search

MCP Server

  • STDIO transport — for MCP harnesses (Zed, opencode, VS Code, Claude Code, etc.)

  • HTTP transport — for remote access and browser clients

  • CORS support — configurable origins for browser-based MCP clients

  • Noise-free startup with UTF-8 encoding (cross-platform, incl. Windows)

CLI Tools

  • crawl — crawl pages from the command line

  • search — search the web via SearXNG

  • crawl-capture — session capture for authenticated crawling

Installation

Docker Compose

The Compose stack includes searxNcrawl + Playwright/Chromium. SearXNG must be provided separately.

cp .env.example .env
# Edit .env: set SEARXNG_URL to your SearXNG instance
docker compose up --build

Variable

Default

Description

MCP_PORT

9555

MCP server HTTP port

LOG_LEVEL

INFO

MCP server log level (DEBUG, INFO, WARNING, ERROR, CRITICAL)

FASTMCP_HTTP_ALLOWED_HOSTS

(FastMCP secure defaults)

JSON list of trusted HTTP Host headers, for example ["mcp.example.com"]

The MCP server is available at http://localhost:9555/mcp.

pip

cd searxNcrawl
python -m venv .venv
source .venv/bin/activate
pip install -e .
playwright install chromium

uv

cd searxNcrawl
uv sync
uv run playwright install chromium

SearXNG (search feature)

The search tool and CLI command require a SearXNG instance with JSON output enabled (search.formats in settings.yml). For all setups you need your own instance — self-hosting is recommended over public instances (rate limits).

Environment variables:

Variable

Example / Recommended

Description

SEARXNG_URL

http://localhost:8888

SearXNG instance URL

SEARXNG_USERNAME

(none)

Optional basic auth user

SEARXNG_PASSWORD

(none)

Optional basic auth pass

SEARCH_RESULT_FIELDS

title,url,content,publishedDate

Comma-separated result fields. Unset = all SearXNG fields. Available: title, url, content, publishedDate, engine, score, category, img_src, thumbnail

Example .env:

SEARXNG_URL=http://localhost:8888
SEARCH_RESULT_FIELDS=title,url,content,publishedDate
LOG_LEVEL=INFO

Config file search order (CLI tools only):

  1. ./.env — current directory

  2. ~/.config/searxncrawl/.env — user config

If no .env exists, .env.example is auto-copied to the user config path.

Usage

MCP Server

Start the server

# STDIO transport (for MCP harnesses)
python -m crawler.mcp_server

# HTTP transport
python -m crawler.mcp_server --transport http --port 8000

# HTTP exposed through a specific public hostname
python -m crawler.mcp_server --transport http --host 0.0.0.0 --allowed-hosts "mcp.example.com"

# HTTP with CORS
python -m crawler.mcp_server --transport http --allowed-hosts "mcp.example.com" --cors-origins "https://app.example.com"

# Docker (HTTP only)
docker compose up --build

MCP client configuration

Python with venv:

{
  "mcpServers": {
    "crawler": {
      "command": "python",
      "args": ["-m", "crawler.mcp_server"],
      "cwd": "/path/to/searxNcrawl",
      "env": { "SEARXNG_URL": "http://your-searxng:8888" }
    }
  }
}

With uv (no manual venv):

{
  "mcpServers": {
    "crawler": {
      "command": "uv",
      "args": ["run", "--directory", "/path/to/searxNcrawl", "python", "-m", "crawler.mcp_server"],
      "env": { "SEARXNG_URL": "http://your-searxng:8888" }
    }
  }
}

Docker (HTTP endpoint):

{
  "mcpServers": {
    "crawler": {
      "url": "http://localhost:9555/mcp"
    }
  }
}

CORS

FastMCP validates the HTTP Host header independently of the address on which the server listens. For remote access, allow the exact externally visible Host header with a comma-separated CLI value:

crawl-mcp --transport http --host 0.0.0.0 --allowed-hosts "mcp.example.com,mcp.internal.example"

Alternatively, use FastMCP's environment setting. It uses JSON-list syntax:

FASTMCP_HTTP_ALLOWED_HOSTS='["mcp.example.com"]' crawl-mcp --transport http --host 0.0.0.0

Browser Origin validation and CORS response headers are separate from Host validation. --cors-origins configures both FastMCP's Origin guard and the CORS middleware using the same normalized, comma-separated values:

crawl-mcp --transport http --cors-origins "http://localhost:3000,https://myapp.com"
crawl-mcp --transport http --cors-origins "*"   # all origins — local dev only

Omitting either allowlist preserves FastMCP's secure defaults (and permits the upstream environment setting to apply). A value of * for Hosts or Origins is an explicit opt-in to broad access and should only be used when that security trade-off is intentional. Without --cors-origins, no CORS headers are sent.

CLI Tools

After pip install -e . (or uv sync), the following commands are available:

# Crawl a page
crawl https://docs.example.com

# Site crawl with depth limit
crawl https://docs.example.com --site --max-depth 2 --max-pages 10 -o docs/

# Clean output (no links)
crawl https://example.com --remove-links

# Search
search "python tutorials"
search "Rezepte" --language de --max-results 5

# Session capture for authenticated crawling
crawl-capture --start-url https://example.com/login \
    --completion-url 'https://example.com/dashboard.*' \
    --output ./state.json

See Session Capture for the full crawl-capture guide.

Python API

from crawler import crawl_page, crawl_page_async, crawl_site, crawl_site_async

# Single page
doc = await crawl_page_async("https://docs.example.com/intro", dedup_mode="exact")
print(doc.markdown)

# Site crawl
result = crawl_site("https://docs.example.com", max_depth=2, max_pages=10)
for doc in result.documents:
    print(f"{doc.status}: {doc.final_url}")

# Authenticated crawl
doc = await crawl_page_async(
    "https://example.com/private",
    auth={"storage_state": "/path/to/state.json"},
)

Reference

  • MCP Tools — full parameter reference for crawl, crawl_site, search

  • Output Formats — Markdown and JSON output structure, including CrawledDocument

  • Session Capture — manual login flow and CDP session export

Configuration

Default config is optimized for documentation sites. Customize via overrides:

from crawler import build_markdown_run_config, RunConfigOverrides

config = build_markdown_run_config(
    RunConfigOverrides(
        delay_before_return_html=1.0,
        mean_delay=1.0,
        scan_full_page=True,
    )
)
doc = await crawl_page_async("https://example.com", config=config)

Dependencies

  • crawl4ai>=0.7.4 — crawler engine

  • playwright>=1.40.0 — browser automation

  • fastmcp>=3.4.3 — MCP server framework

  • httpx>=0.27.0 — HTTP client for SearXNG

  • tldextract>=5.1.2 — domain parsing for site crawls

License

MIT — © 2026 DDM – Das Digitale Momentum GmbH & Co KG

Available Tools

3 tools
crawlB

Crawl one or more URLs and extract content as markdown or JSON.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlsYesList of URLs to crawl (required). Accepts a single URL or multiple URLs.
timeoutNoPer-URL timeout in seconds (default: 15, must be >= 1)
dedup_modeNoMarkdown dedup mode - "exact" (default) or "off"exact
concurrencyNoMaximum concurrent crawls (default: 3)
remove_linksNoRemove all links from the markdown output (default: false)
output_formatNo'markdown' (default) or 'json' - markdown: Clean concatenated markdown with URL headers and timestamps - json: Full JSON with metadata, references, and statisticsmarkdown
storage_stateNoPath to Playwright storage_state JSON for authenticated crawling

TDQS

B3.1/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

No annotations are provided, so the description carries the full burden of behavioral disclosure. It only states what the tool does functionally, but does not mention side effects, authentication needs, rate limits, or any limitations. For a crawler that might hit external sites, this transparency gap is significant.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that efficiently conveys the core purpose. There is no wasted language, and it is appropriately front-loaded.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 7 parameters and no output schema, but the input schema is thorough. The description is minimal, yet combined with the schema it provides enough to invoke the tool. However, it lacks any context about output structure, error behavior, or when to choose this over crawl_site, leaving some gaps for an agent deciding on usage.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The schema has 100% description coverage with detailed parameter meanings, so the description does not need to repeat them. The tool description adds minimal value beyond the schema, merely echoing the markdown/JSON choice already present in output_format. Baseline 3 is appropriate as the schema does the heavy lifting.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description clearly states a specific verb+resource: 'Crawl one or more URLs and extract content as markdown or JSON.' This distinguishes it from search and is reasonably specific. However, it does not explicitly differentiate from the sibling tool crawl_site, so it misses the full differentiation.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines2/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description provides no guidance on when to use this tool vs alternatives. It does not mention appropriate use cases, exclusions, or when to prefer crawl_site or search. The extensive parameter descriptions in the schema do not compensate for this lack of contextual usage guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

crawl_siteB

Crawl an entire website starting from a seed URL using DFS strategy.

ParametersJSON Schema
NameRequiredDescriptionDefault
urlYesThe seed URL to start crawling from
timeoutNoOverall site crawl timeout in seconds (default: 120, must be >= 1)
max_depthNoMaximum depth to crawl (default: 2, 0 = seed page only)
max_pagesNoMaximum number of pages to crawl (default: 25)
dedup_modeNoMarkdown dedup mode - "exact" (default) or "off"exact
remove_linksNoRemove all links from the markdown output (default: false)
output_formatNoOutput format - "markdown" (default) or "json" - markdown: Clean concatenated markdown with URL headers and timestamps - json: Full JSON with metadata, references, and crawl statisticsmarkdown
storage_stateNoPath to Playwright storage_state JSON for authenticated crawling
include_subdomainsNoWhether to include subdomains in the crawl (default: false)

TDQS

B3.2/5.0
Behavior2/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden of behavioral disclosure, but it only mentions the DFS traversal strategy. It does not disclose potential impacts (e.g., rate limits, resource consumption), authentication requirements, the meaning of 'entire website' (same-domain vs. subdomains), or what the output will look like. This is insufficient for a crawler that can visit many pages.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is a single, concise sentence that front-loads the tool's purpose and key strategy. Every word adds value, and there is no fluff or repetition of schema information.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness2/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

The tool has 9 parameters, no output schema, and no annotations, yet the description is only one line. It does not explain what the return value looks like, any limits or side effects, or how the parameters relate to the crawling process. For a complex tool like this, the description is notably incomplete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

The input schema already provides detailed descriptions for all 9 parameters, covering 100% of the schema. The description adds no additional parameter semantics, so the baseline score of 3 applies — it neither enhances nor detracts from the schema's clarity.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description uses a specific verb ('crawl'), a concrete resource ('entire website'), and a strategy ('DFS'), which clearly states what the tool does. However, it does not distinguish this tool from the sibling tool named 'crawl', leaving ambiguity about how they differ.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies the tool is for crawling entire websites via DFS, giving some context. But it does not explicitly state when to use this tool over the 'crawl' or 'search' sibling tools, and it offers no exclusions or alternative guidance. The presence of a sibling named 'crawl' makes this gap more noticeable.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Tool Schema Changelog

Recent tool additions, removals, and schema changes observed during successful MCP inspections.

  1. 3 tool updatesv0.30.0
    • First observedcrawl
    • First observedcrawl_site
    • First observedsearch

TDQS

A3.5/5.0

Scored across 3 tools

Disambiguation5/5

Each tool has a distinct purpose: 'search' queries the web, 'crawl' extracts specific URLs, and 'crawl_site' handles entire websites. There is no ambiguity between them.

Naming Consistency4/5

Tool names use a clear verb-based pattern ('crawl', 'search', 'crawl_site'), though 'crawl_site' mixes a verb with an object while the others are single verbs. The pattern is still predictable and readable.

Tool Count5/5

With only 3 tools, the server is well-scoped for its combined search and crawl functionality. Each tool addresses a core need without unnecessary bloat.

Completeness5/5

The tool surface covers the full lifecycle of its domain: searching the web, crawling individual URLs, and crawling entire sites. No obvious gaps or dead ends.

Maintenance

ActivityNo data
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers

  • A
    license
    Not graded
    quality
    C
    maintenance
    MCP server for web search and content extraction using DuckDuckGo or SearXNG, with Playwright-based fetching and LLM-powered data extraction.
    139
    MIT
  • A
    license
    Not graded
    quality
    D
    maintenance
    MCP server for internet search via direct Google and DuckDuckGo HTML scraping with AI-powered result normalization and optional summarization, requiring no API keys for search.
    MIT
  • A
    license
    A
    quality
    D
    maintenance
    A privacy-focused web search and content extraction MCP server. It integrates SearxNG with fallback to Google scraping, featuring relevance ranking, security-aware search, and rate limiting.
    3
    MIT
  • A
    license
    A
    quality
    B
    maintenance
    MCP server for local web search via SearXNG, providing unlimited queries without API keys or cost, with automatic fallback to public instances.
    3
    1
    MIT