searxng-rank-mcp
Used as the default LLM backend for the orchestrator's iterative search/evaluate loop, calling a local Ollama model (e.g. llama3.1:8b) through its OpenAI-compatible API to judge coverage, refine queries, and synthesize cited answers.
The orchestrator can call any OpenAI-compatible LLM endpoint, including OpenAI's API, to drive the iterative search/evaluate loop and synthesize grounded answers with [n] citations.
Serves as the core retrieval backend: the server queries any reachable self-hosted SearXNG metasearch instance to fetch multi-engine web results, then extracts, reranks, dedups, and returns them with citation markers.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@searxng-rank-mcpsearch for recent solid-state battery advances and rank the best sources"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
searxng-rank
Self-hosted web search that gets you closer to "Perplexity-quality" answers without a proprietary index. It runs the classic research loop — retrieve, extract, rerank, evaluate, iterate — over any reachable SearXNG instance, and exposes the result two different ways:
An orchestrator service that owns the iterative search/evaluate loop and calls any OpenAI-compatible LLM directly.
A thin, LLM-free MCP tool surface (
search/fetch/rerank) for agent hosts (Open WebUI, OpenCode, Claude, a raw vLLM + tool-calling model) that want to drive their own loop.
The two surfaces share one deterministic core: SearXNG retrieval → page-text extraction (Trafilatura) → cross-encoder rerank (FlashRank) → dedup, plus bounded "widen" retries when results look weak.
Inspiration, not affiliation. The shape of this tool is inspired by the Exa search API (a search call that returns ranked results plus a grounded, cited answer), and its architecture draws on Vane (the open-source Perplexity).
searxng-rankis our own name for reimplementing that contract self-hosted over SearXNG — it is not affiliated with, or a reference to, Exa.
Why SearXNG?
SearXNG is a self-hostable metasearch aggregator over hundreds of engines. Pointing the pipeline at "any reachable SearXNG" means no API keys, no per-query cost, and a stable endpoint you control. The project degrades gracefully: if SearXNG is unreachable it tells you so instead of crashing, and if the reranker model is missing it falls back to SearXNG's own ordering.
Related MCP server: WebFetch.MCP
Quickstart
Docker (recommended)
docker compose up -d brings up the whole stack together:
SearXNG (
searxng/searxng:latest) on:8888, configured fromsettings/searxng.yml(JSON format on, limiter off).the MCP server (
searxng-rank-mcp) on:8000.Qdrant on
:6333(vector store; the Phase-2 cache).
The MCP server finds SearXNG automatically by compose service name, so nothing
else to configure. Then point a host at
http://<host>:8000/mcp/ (see MCP integration).
Dev (uv)
uv sync # install dependencies
uv run searxng-rank "your query" # run the orchestrator end-to-end
uv run searxng-rank-mcp # run the MCP tool server (:8000)
uv run pytest # tests
uv run ruff check . # lintRun a SearXNG separately (bundled image or your own) and set
SEARXNG_BASE_URL=http://<host>:<port> for the MCP to find it — discovery
tries an explicit URL first, then the same-network service name, then common
local defaults (see below).
How it works
User query
→ (any reachable) SearXNG retrieve, multi-engine
→ Trafilatura extract clean page text
→ FlashRank cross-encoder rerank (quality threshold ~0.7)
→ dedup + bounded widen retry / widen when results look weak
→ [orchestrator] LLM evaluate "is this enough? what's the next query?"
→ loop (refine) or synthesize answer with [n] citationsThe deterministic part (everything above the LLM) is shared by both surfaces and is LLM-free. The semantic part (is this enough? what's the next query? what's the answer?) lives in the host LLM — the orchestrator calls it directly; the MCP surface lets your host model own it.
Results carry a stable [n] Title — URL marker plus a short Signals: block
(score stats, threshold, domain diversity) so an LLM can judge coverage and
cite sources as [n] / markdown links.
Configuration
All configuration is via environment variables (see
.env.example):
Variable | Purpose | Default |
| Explicit SearXNG endpoint (highest precedence) | — |
| Same-network service to probe |
|
| SearXNG instance secret (compose only) | — |
| MCP bind address |
|
| LLM for the orchestrator (OpenAI-compatible) | Ollama / |
| Vector store |
|
| Cross-encoder model |
|
| Orchestrator loop budget |
|
| Rerank pass threshold |
|
SearXNG discovery (see src/search/discovery.py): an explicit
SEARXNG_BASE_URL wins; otherwise the same-network service name is tried;
then common local defaults. Every candidate is probed with a cheap
GET /search?format=json; the first that answers is used. A healthy result
is cached and sticky; an unhealthy one is re-probed on the next call, so a
late-starting SearXNG recovers without a restart.
MCP integration
The MCP surface is Streamable HTTP (the transport Open WebUI supports natively). One server hosts many clients; the host model drives the loop and cites results. See docs/mcp-integration.md for per-host setup (Open WebUI, OpenCode, raw vLLM) and how citations render.
Project layout
searxng-rank/
├── src/
│ ├── main.py CLI entry point (orchestrator)
│ ├── pipeline.py LLM-free SearchService (search/fetch/rerank)
│ ├── models.py shared dataclasses
│ ├── search/ SearXNG client + endpoint discovery
│ ├── extraction/ Trafilatura page-text extraction
│ ├── ranking/ FlashRank cross-encoder reranker
│ ├── orchestration/ LLM client + iterative search/evaluate loop
│ └── mcp/ Streamable-HTTP MCP tool surface
├── tests/ unit + in-memory-MCP + re-probe tests
├── settings/searxng.yml SearXNG config (JSON on, limiter off)
├── docker-compose.yml SearXNG + MCP + Qdrant
├── Dockerfile the MCP server image
└── docs/ architecture, MCP integration, prior art, embeddingStatus
A working Phase-1 search: any-reachable-SearXNG provider, cross-encoder rerank with quality threshold, iterative orchestrator loop, and the LLM-free MCP tool surface. The persistent collaborative content/embedding cache (Phase 4) is design-only — see docs/embedding.md and docs/prior-art.md.
License
MIT — see LICENSE.
This server cannot be deployed
Maintenance
Related MCP Connectors
LLM-ready web search + instant answers + URL-to-clean-text fetch for agents and RAG.
Web search for AI agents — one tool across 6 engines, routed to the cheapest + cached.
Web search, fetch, extract, and research for AI agents. Markdown output + AI-synthesized answers.
Multi-engine search for AI agents. Trust scoring, local corpus, MCP-native. Self-hostable, BYOK.
Related MCP Servers
- AlicenseAqualityAmaintenanceWeb search (embedded SearXNG), content extraction, and library docs indexing with hybrid search. No API keys required.61,488 PyPI18Apache 2.0
- AlicenseAqualityCmaintenanceEnables local LLMs to search the web and fetch clean content from URLs without API keys, using SearxNG and Mozilla Readability.238MIT
- FlicenseNot gradedqualityDmaintenanceProvides free web search, content fetching, image search, and deep research via SearXNG, no API keys required.-
- FlicenseNot gradedqualityAmaintenanceEnables AI agents to perform web searches and page fetching/crawling through self-hosted SearXNG and Crawl4AI, without depending on commercial search or scraping APIs.-