Skip to main content
Glama
ptrken01

searxng-mcp-scraper

by ptrken01
README.md
# searxng-mcp-scraper

> MCP server exposing SearXNG search and HTTP fetch tools for AI agents.

A self-contained [Model Context Protocol](https://modelcontextprotocol.io) server that turns
[SearXNG](https://docs.searxng.org/) (a self-hosted metasearch engine) into two MCP tools
you can wire into Claude Desktop, Cursor, or any MCP-compatible client:

- **`search(query, ...)`** — ranked web results from SearXNG's JSON API
- **`fetch(url)`** — cleaned main text of any public http(s) URL

No API keys, no third-party tracking, no rate limits beyond what your SearXNG instance imposes.

## Installation

```bash
# From source (recommended during dev)
git clone https://github.com/your-org/searxng-mcp-scraper
cd searxng-mcp-scraper
uv sync --extra dev
uv pip install -e .

# Or with plain pip
pip install -e .
```

## Configuration

The server reads configuration from environment variables (or a `.env` file at startup):

| Variable | Required | Default | Description |
|----------|----------|---------|-------------|
| `SEARXNG_BASE_URL` | **yes** | — | Base URL of your SearXNG instance, e.g. `http://localhost:8888` |
| `SEARXNG_DEFAULT_CATEGORY` | no | `general` | Default search category |
| `SEARXNG_DEFAULT_ENGINES` | no | `wiby,naver,seznam,marginalia,wikipedia,duckduckgo_html` | Default engine list when a caller doesn't pass `engines`. Curated to the keyless engines that return results from this machine's egress IP (Brave/Google/Startpage/DuckDuckGo are rate-limited here). Set to empty to use SearXNG's own default set. |
| `SEARXNG_DEFAULT_LANGUAGE` | no | `en` | Default language code |
| `SEARXNG_DEFAULT_SAFESEARCH` | no | `0` | Default safesearch: `0`/`1`/`2` |
| `SEARCH_MAX_RESULTS_CAP` | no | `50` | Hard cap on results per call |
| `SEARCH_TIMEOUT_S` | no | `30` | HTTP timeout (seconds) for search |
| `FETCH_TIMEOUT_S` | no | `20` | HTTP timeout (seconds) for fetch |
| `FETCH_MAX_BYTES` | no | `5_000_000` | Refuse to read more than this many bytes per fetch |
| `FETCH_MAX_REDIRECTS` | no | `5` | Max HTTP redirects to follow |
| `FETCH_ALLOW_PRIVATE` | no | `false` | Allow fetching private/loopback URLs (SSRF guard, default on) |
| `USER_AGENT` | no | `searxng-mcp-scraper/0.1` | User-Agent header for all requests |
| `LOG_LEVEL` | no | `INFO` | Stderr log verbosity: `DEBUG`/`INFO`/`WARNING`/`ERROR` |
| `MCP_HOST` | no | `127.0.0.1` | Bind host when `--transport streamable-http` |
| `MCP_PORT` | no | `8765` | Bind port when `--transport streamable-http` |
| `MCP_CORS_ORIGINS` | no | `http://localhost,http://localhost:*,http://127.0.0.1,http://127.0.0.1:*` | Comma-separated CORS allow-origins. Use `*` to allow any origin (not recommended on LAN). |
| `BLOG_SCRAPE_OUTPUT_DIR` | no | `~/scraped-blogs` | Directory for `scrape_blog` / `deep_scrape` Markdown output and assets |
| `BLOG_SCRAPE_MAX_POSTS` | no | `200` | Max post pages per blog scrape |
| `BLOG_SCRAPE_CONCURRENCY` | no | `5` | Max parallel post fetches |
| `BLOG_SCRAPE_DISCOVERY_PAGES` | no | `10` | Max HTML index pages to scan when no feed is found |
| `DEEP_SCRAPE_MAX_DOCUMENTS_PER_BLOG` | no | `100` | Max linked documents downloaded/extracted by one `deep_scrape` call |
| `DEEP_SCRAPE_MAX_DOCUMENT_BYTES` | no | `50_000_000` | Refuse a single linked document above this size |
| `DEEP_SCRAPE_MAX_IMAGES_PER_BLOG` | no | `200` | Max images downloaded by one `deep_scrape` call |
| `DEEP_SCRAPE_MAX_IMAGE_BYTES` | no | `15_000_000` | Refuse a single image above this size |
| `DEEP_SCRAPE_CONCURRENCY` | no | `5` | Max parallel document/image downloads during `deep_scrape` |
| `FIRECRAWL_FALLBACK_ENABLED` | no | `true` | Retry a failed/thin static fetch through Firecrawl (see below) |
| `FIRECRAWL_BASE_URL` | no | `http://127.0.0.1:8788` | Firecrawl API base — the local keyless proxy by default |
| `FIRECRAWL_API_KEY` | no | — | Optional. Empty uses the keyless free tier |
| `FIRECRAWL_TIMEOUT_S` | no | `60` | Timeout for a fallback scrape (Firecrawl renders before responding) |
| `FIRECRAWL_MIN_TEXT_CHARS` | no | `500` | Static extractions shorter than this are treated as failures and retried |

> **Note:** your SearXNG instance must have `json` enabled in `search.formats` under `settings.yml`,
> or `search` will return `searxng_unavailable` with a 403.

### PDF extraction (optional, local, keyless)

`deep_scrape` downloads linked PDFs and extracts their text. By default it
uses `markitdown`. For a **much faster and cleaner** text extraction on
text-based PDFs, install the optional extra:

```bash
uv sync --extra pdf-extras          # or: pip install "searxng-mcp-scraper[pdf-extras]"
```

This pulls in [`pdf-inspector`](https://github.com/firecrawl/pdf-inspector) —
Firecrawl's **open-source Rust PDF engine** (no API key, no cloud, runs
locally). When installed, PDF extraction routes through it first:

- **text-based PDFs** are read straight from the PDF internals (fonts, text
  operators) in milliseconds.
- **scanned / image-only PDFs** are correctly *classified* as needing OCR
  (the library reports `pdf_type: scanned`). It has no OCR of its own, so
  those fall through to markitdown (which may have an OCR backend); if
  markitdown also returns nothing, the document surfaces an honest
  `DocumentExtractionFailed` instead of a silent empty string.

The dependency is optional: without it, PDFs simply use markitdown. Nothing
in the core package depends on a native wheel.

### Firecrawl fallback (JS-rendered pages)

`fetch` does a plain HTTP GET and runs trafilatura over the HTML. That is fast and
free, but it returns nothing useful for client-rendered pages: the served body is an
empty root div, so trafilatura scrapes up the meta description and a nav label and
returns a couple hundred characters of noise. Measured on real sites:

| URL | static | with fallback |
|-----|--------|---------------|
| `vercel.com/templates` | 208 chars | **15,972** |
| `notion.so/product` | 156 chars | **7,737** |
| `linear.app/method` | 326 chars | **1,126** |
| `excalidraw.com` | 57 chars | **655** |
| `app.slack.com` | 2,764 chars | 2,764 (static kept — already good) |

When a fetch errors, returns empty text, or returns fewer than
`FIRECRAWL_MIN_TEXT_CHARS`, the URL is retried once through Firecrawl's
`/v2/scrape`, which renders the page server-side and returns clean markdown.
Results carry `extractor: "firecrawl"` so you can tell which path produced them.

Firecrawl's hosted API has a **keyless free tier** for `/v2/scrape` and `/v2/search`,
so this costs nothing and needs no account. The catch: the tier rejects *any*
`Authorization` header, while most clients insist on sending one. `FIRECRAWL_BASE_URL`
therefore points at a small local proxy that strips the header:

```bash
python3 scripts/firecrawl_keyless_proxy.py   # 127.0.0.1:8788
```

The proxy is vendored here (no external dependency) and forwards a real key
if you set `FIRECRAWL_UPSTREAM_KEY` in its environment — which unlocks
`/v2/map`, `/v2/crawl`, and `/v2/extract` through the same URL. Key-gated
paths without a key get a clear error rather than a confusing upstream 403.

Set `FIRECRAWL_API_KEY` and point `FIRECRAWL_BASE_URL` at `https://api.firecrawl.dev`
to skip the proxy and use a real account instead.

**The fallback never overrides a guard.** `invalid_url` and `private_network_blocked`
are SSRF/scheme decisions, not extraction failures — routing those through a
third-party renderer would defeat the guard, so they are excluded. `oversize` is
excluded too (re-rendering the same huge page just burns the cap again). If the
fallback fails or returns *less* text than the static path, the original result is
kept unchanged. Set `FIRECRAWL_FALLBACK_ENABLED=false` for fully air-gapped runs.

## Running

The server supports two transports. Pick the one your MCP client speaks:

### stdio (default — Claude Desktop, Cursor, etc.)

```bash
# From a checkout, with .env in the cwd
SEARXNG_BASE_URL=http://localhost:8888 searxng-mcp-scraper

# Or via uv
SEARXNG_BASE_URL=http://localhost:8888 uv run searxng-mcp-scraper

# Or as a module
SEARXNG_BASE_URL=http://localhost:8888 python -m searxng_mcp_scraper
```

The server speaks MCP over stdio. It writes logs to stderr; stdout is reserved for the
JSON-RPC stream and must not be polluted.

### streamable-http (llama-ui, Open WebUI, browser-based clients)

```bash
SEARXNG_BASE_URL=http://localhost:8888 uv run searxng-mcp-scraper --transport streamable-http
```

This binds to `http://127.0.0.1:8765/mcp` by default. Override with `MCP_HOST` / `MCP_PORT`
env vars. CORS is pre-configured for `http://localhost:*` and `http://127.0.0.1:*` — any
browser on your machine can connect.

If you need to bind to `0.0.0.0` (LAN), set `MCP_CORS_ORIGINS` to a narrower list. There is
no auth in v1 — do not expose this to the public internet without putting it behind a
reverse proxy with auth.

## llama-ui / Open WebUI config

For browser-based MCP clients, run the server in HTTP mode and paste the URL into
the client's "MCP server URL" field:

1. In one terminal:
   ```bash
   SEARXNG_BASE_URL=http://localhost:8888 uv run searxng-mcp-scraper --transport streamable-http
   ```
2. In the UI, add a new MCP server with URL `http://127.0.0.1:8765/mcp` (no auth
   headers required for local).
3. The `search` and `fetch` tools will appear in the model's tool list.

## Claude Desktop config

Add this to `~/Library/Application Support/Claude/claude_desktop_config.json`:

```json
{
  "mcpServers": {
    "searxng-scraper": {
      "command": "searxng-mcp-scraper",
      "env": {
        "SEARXNG_BASE_URL": "http://localhost:8888"
      }
    }
  }
}
```

If you installed from source and the binary isn't on your `PATH`, point `command` at `uv`:

```json
{
  "mcpServers": {
    "searxng-scraper": {
      "command": "uv",
      "args": ["--directory", "/absolute/path/to/searxng-mcp-scraper", "run", "searxng-mcp-scraper"],
      "env": {
        "SEARXNG_BASE_URL": "http://localhost:8888"
      }
    }
  }
}
```

Restart Claude Desktop; you should see `search` and `fetch` tools appear.

## Tool reference

### `search(query, ...) -> {results, suggestions, number_of_results}`

| Parameter | Type | Default | Description |
|-----------|------|---------|-------------|
| `query` | string | (required) | Search query string |
| `categories` | list[string] | `["general"]` | SearXNG categories, e.g. `["general"]`, `["images"]` |
| `engines` | list[string] | `null` | Restrict to specific engines, e.g. `["google", "bing"]` |
| `language` | string | `"en"` | Language code |
| `pageno` | int | `1` | Page number (1-indexed) |
| `time_range` | `"day"`/`"month"`/`"year"`/`null` | `null` | Time filter |
| `safesearch` | `0`/`1`/`2` | `0` | Safe-search level |
| `max_results` | int | `10` | Cap on results returned (1–50) |

### `fetch(url) -> {url, final_url, content_type, text, byte_count}`

| Parameter | Type | Description |
|-----------|------|-------------|
| `url` | string | An `http://` or `https://` URL |

Strips `<script>`, `<style>`, `<nav>`, `<header>`, `<footer>`, `<aside>` blocks via
`trafilatura`, and collapses whitespace. Refuses non-http(s) schemes and (by default) private
network targets. Returns a structured error on timeout, oversize, or non-2xx — never crashes.

### `scrape_blog(blog_url) -> {output_path, post_count, ...}`

Discovers a blog's RSS/Atom feed first, falls back to scanning HTML index pages, fetches post
text, and writes one Markdown file under `BLOG_SCRAPE_OUTPUT_DIR`.

### `deep_scrape(blog_url) -> {output_path, assets_dir, documents_*, images_*}`

Does everything `scrape_blog` does, plus:

- discovers linked documents from each post (`.pdf`, `.docx`, `.csv`, `.md`, etc.)
- downloads raw document files into `deep_assets/`
- extracts document text with MarkItDown and embeds it in the Markdown output
- discovers images from `<img src>`, lazy-load attributes, `srcset`, and direct image links
- downloads bounded image files into `deep_assets/images/`
- records per-image URL, saved path, content type, byte count, and any failure in the Markdown

The returned summary includes `images_found`, `images_saved`, and `image_byte_count`. One bad
image/document is recorded in-band and does not fail the whole blog scrape.

## Development

```bash
uv run pytest          # run the full test suite
uv run mypy src        # static type checks
uv run ruff check      # lint
```

## License

MIT.

TDQS

A4/5.0

Scored across 4 tools

Disambiguation4/5

The tools are mostly distinct: search performs queries, fetch retrieves a single URL, while scrape_blog and deep_scrape both focus on blog content aggregation. The overlap between scrape_blog and deep_scrape is clear (deep_scrape is an extended version), but both descriptions are detailed enough to avoid confusion in practice.

Naming Consistency3/5

The names mix conventions: 'search' and 'fetch' are single generic verbs, 'scrape_blog' follows a verb_noun pattern, and 'deep_scrape' is an adjective_verb compound. While each name is readable and descriptive, the lack of a uniform pattern makes the set feel slightly inconsistent.

Tool Count5/5

With only 4 tools, the server is well-scoped for its purpose. Each tool covers a distinct aspect of web scraping and search: querying, single-page extraction, blog aggregation, and deep scraping with document extraction. No tool feels redundant or unnecessary.

Completeness4/5

The tool surface covers the core workflows: web search, URL fetching, blog scraping, and extended scraping with document extraction. Minor gaps exist, such as a generic crawling tool for non-blog sites, but these are not critical given the server's stated focus on search and blog content.

Maintenance

ActivitySlowing
ResponsivenessNo issues