research-mcp
Allows searching the web using a SearXNG instance, providing search results with titles, URLs, and snippets.
Click on "Deploy Server".
Wait a few minutes for the server to deploy. Once ready, it will show a "Started" state.
In the chat, type
@followed by the MCP server name and your instructions, e.g., "@research-mcpsearch for MCP server examples and read the top result"
That's it! The server will respond to your query, and you can continue using it as needed.
Here is a step-by-step guide with screenshots.
research-mcp
A stateless MCP facade that hides a pyramid of search/read providers behind a single streamable-http MCP endpoint and exposes just 4 clean tools with good help texts. An LLM gets a simple "search → read" toolset (or both at once); behind it, several providers are tried, merged, and failed over automatically.
The app does no authentication — it is published through Traefik + basicAuth
on the host. It holds no application state: the only thing persisted is a log
file under data/ (kept on a volume).
Works with zero keys
make run on an empty .env is enough — no key, no SearXNG, nothing to deploy
first. Two instances need no configuration at all and are therefore always
enabled: duckduckgo (search — the no-JS html.duckduckgo.com SERP, one
request per query, no token handshake) and trafilatura (read — local
HTML→Markdown extraction). Together they are the floor: search and read both
work out of the box.
Everything else is an upgrade on top of that floor. A self-hosted SearXNG
(SEARXNG_URL) is optional; when it is configured it sits ahead of
DuckDuckGo in the pipeline, so SearXNG keeps its copy of every url both of them
return and DuckDuckGo only adds what nobody ahead of it had. (With the reranker
on — the default once JINA_API_KEY is set — the merged list is reordered by
relevance across all sources before the trim, so the final ordering can still
change.) Every paid vendor lights up the moment its key appears.
The floor is deliberately modest: DuckDuckGo is a scraped SERP, not an API, so it paces itself to one query per 45s (skipping, never waiting, when the slot is taken) and reports a block or a captcha as a failure rather than as an empty result set. That pace is the one measured for this upstream: DuckDuckGo blocks by IP for 7-8 minutes after a burst, and a block taken here would also silence a SearXNG that reaches DuckDuckGo over the same address.
Related MCP server: Web Search MCP
Tools
Tool | What it does |
| Search across all enabled providers, merge + dedup → ranked list (title, URL, snippet). Search only. |
| One page or PDF → clean Markdown. Auto-detects type, walks the read pipeline (light → heavy) until one succeeds. |
| Up to 20 urls concurrently → |
| Search and read the top hits in one call → |
The tool descriptions cross-reference each other (when to take this one, when to
take another), so the model gets a routing graph instead of four independent
texts. search_and_read is the default for research; web_search is for links
only; read_pages / read_page are for urls that are already known.
The batch tools (read_pages, search_and_read) cap each page at
READ_BATCH_MAX_CHARS and mark the cut with [содержимое обрезано на N символах]; read_page always returns the page in full.
Architecture: types + instances
Providers are plugins. We separate:
type — an implementation class (e.g. the
searxngsearch provider), one per module insrc/providers/, registered with@register("type").instance — a configured copy of a type with its secrets/URL resolved from named environment variables (multiple instances of one type are allowed, e.g.
tavily-1/tavily-2with different keys).
Which instances exist and the order each pipeline tries them is configured in
code (src/pipeline_config.py); keys/URLs come from ENV by variable name.
Search pipeline (
searxng → duckduckgo → brave → tavily-search → firecrawl-search → jina-search → xmlriver → parallel → octen → linkup → youcom → serper → exa): enabled instances run concurrently; results are merged and deduplicated by normalized URL (earlier pipeline position wins). Position is therefore a dedup preference, not a cost gate — every enabled instance is called on every query, so cost scales with how many keys are set. WhenJINA_API_KEYis set (andSEARCH_RERANK_ENABLEDis not turned off), the full merged list is then reranked byjina-reranker-v3.5so the trim tonum_resultskeeps the most relevant hits instead of a blind pipeline-order prefix; any rerank failure falls back to the merge order.searxng,duckduckgoandbraveadditionally throttle themselves locally (one query per 45s, 45s and 1.1s respectively, each matching a measured upstream limit — DuckDuckGo shares SearXNG's, since both reach the same engine from the same address); when the slot is taken they skip the current search instead of waiting for it.duckduckgois the only search instance that needs no configuration, which is why it sits directly behindsearxng: it is on everywhere, but a deployment that runs SearXNG must keep SearXNG's copy of every shared url.Read pipeline (
youtube → instagram → instagram-profile → trafilatura → jina → crawl4ai → tavily-1 → tavily-2 → firecrawl → brightdata): here the order IS a cost gate — it stops at the first sufficient answer, andbrightdata(the anti-bot unlocker) sits last so it only ever sees pages everything cheaper already bounced off. The first three are url-specific readers: each is offered only the urls it recognises, ahead of the probe, and its answer is final (see below). A single probe GET classifies the url. PDFs (Content-Type /.pdf/%PDFmagic) are extracted with pypdf — except a PDF with no text layer (a scan), which falls through into the chain so the remote readers get a shot at it with their own parsers, with pypdf's notice kept as the last resort. jina's OCR tier joins that attempt only when the url path ends in.pdfand jina is keyed; for HTML, that same body is handed totrafilaturaso the hot path never GETs twice, then the remaining instances are tried in order and the first to return content>= FALLBACK_MIN_CHARSwins. A YouTube video url (youtube.com/watch?v=…,/shorts/…,/live/…,/embed/…,youtu.be/…) is answered before the probe with the video's transcript — title, channel, description and the captions as timestamped paragraphs — fetched from YouTube's own player API (the unofficial ANDROID client, the same oneyoutube-transcript-apiuses). The track in the spoken language wins — a manual one when it exists, the auto-generated one otherwise. The spoken language comes from the audio track marked "original" on an auto-dubbed video (which carries auto-generated captions for every dub), and from the single auto-generated track on a plain one. A video without captions, or a fetch that fails, falls through to the normal chain above. An Instagram video url (instagram.com/reel/…,/reels/…,/p/…,/tv/…, and/<username>/reel/…,/<username>/p/…) is answered the same way, with a transcript of its audio — author, caption and the speech as timestamped paragraphs. The post comes from one anonymous request to Instagram's web GraphQL API (the logged-out queryyt-dlpuses); Groq Whisper (whisper-large-v3-turbo) fetches the post's audio-only DASH track by url and transcribes it, so the server never downloads the media itself. Enabled byGROQ_API_KEY; without it, or when the post is private, login-gated or has no video, the url falls through to the normal chain. An Instagram profile url (instagram.com/<username>/) is answered with the profile's posts, 12 at a time, from the same anonymous GraphQL API (the logged-out profile posts query): per post its date, kind (reel / video / carousel / photo), link and caption; a reel or video link can be read again for its transcript (withGROQ_API_KEY). A last lineNext page: …/<username>/?after=<cursor>points to the next 12 — reading that url continues the list. Needs no key. Hashtag pages are not covered: Instagram serves them only to a logged-in account.
Cross-cutting: one transient retry (5xx / transport errors) with a short backoff;
402 (out of credits) / 429 (rate limited) are treated as a provider failure →
next instance (this is what makes tavily-1 → tavily-2 fail over). Vendors
that report an empty balance with some other 4xx — serper answers 400 {"message": "Not enough credits"}, octen 403 "Insufficient balance" — are recognised by
the body and logged as out of credits too, so an unpaid account never reads as
a broken API.
An instance is enabled only if its required env var(s) are set; otherwise it
is skipped with a log line. duckduckgo and trafilatura need no config (always
on); jina works keyless (its key is optional). At startup the server requires at
least one search and one read instance — a condition those two always satisfy, so
the check now only catches a broken pipeline_config.py.
Adding a provider
Write
src/providers/<type>.pywith a class decorated@register("<type>")implementingSearchProvider.search(...)orReadProvider.read(...).Import the module in
src/providers/__init__.py(so the decorator runs).Add an
Instance("name", "<type>", api_key_env="YOUR_ENV_NAME")line insrc/pipeline_config.pyand reference itsnameinSEARCH_PIPELINE/READ_PIPELINE. Use the ENV var NAME, never a value.Document the env var in
.env.example.
Quick start
make install # create .venv + install dev/test deps
cp .env.example .env # fill in the keys you have (shortcut: make env)
make test # run tests
make run # run the server (streamable-http on MCP_HOST:MCP_PORT, endpoint /mcp)Configuration
All config comes from ENV / .env (see .env.example). Provider secrets/URLs
are read by name in the instance loader, not declared as Settings fields. The
non-secret knobs (all defaulted): MCP_HOST, MCP_PORT, LOG_LEVEL,
LOG_FILE, LOG_ROTATION, LOG_RETENTION, REQUEST_TIMEOUT,
FALLBACK_MIN_CHARS, READ_PAGES_CONCURRENCY, READ_BATCH_MAX_CHARS (per-page
content budget of the batch tools — read_pages and search_and_read; beyond it
the markdown is cut and marked, read_page is never truncated, 0 disables the
cut), RETRIES,
SEARCH_RERANK_ENABLED, JINA_TOKEN_BUDGET, ALLOW_PRIVATE_NETWORK (escape
hatch for the SSRF guard: true lets read_page fetch private/loopback
addresses, which are blocked by default). The read_pages
per-call url cap is a fixed 20 (hard constant, matching the tool description) —
not configurable.
Provider env vars — duckduckgo and trafilatura take none and are always on;
everything below is optional on top of them: SEARXNG_URL, BRAVE_API_KEY,
SERPER_API_KEY, EXA_API_KEY, JINA_API_KEY
(one key enables the jina reader in keyed mode, the jina-search provider and
the search reranker; the reader alone also works keyless), CRAWL4AI_URL +
CRAWL4AI_TOKEN, TAVILY_1_API_KEY, TAVILY_2_API_KEY, FIRECRAWL_API_KEY.
The Tavily and Firecrawl keys each enable two instances — the reader and the
search provider — because both vendors sell search and extract off one key, out
of one shared monthly pool. Search runs on every query and will drain that pool
well before the readers do; when it runs out, both halves stop working.
Keyless until registered: XMLRIVER_USER_ID + XMLRIVER_API_KEY (Yandex SERP),
PARALLEL_API_KEY, OCTEN_API_KEY, LINKUP_API_KEY, YOUCOM_API_KEY, and
BRIGHTDATA_API_KEY + BRIGHTDATA_ZONE. GROQ_API_KEY enables the instagram
reader (Instagram transcripts); the youtube and instagram-profile readers
take no key and are always on.
Proxy
Any external instance can be routed through its own SOCKS5/HTTP proxy by
setting <INSTANCE>_PROXY — useful for clean egress past IP-based blocks (e.g.
Cloudflare in front of Exa). Supported per instance: EXA_PROXY, BRAVE_PROXY, SERPER_PROXY,
JINA_PROXY, TAVILY_1_PROXY, TAVILY_2_PROXY, FIRECRAWL_PROXY,
XMLRIVER_PROXY, PARALLEL_PROXY, OCTEN_PROXY, LINKUP_PROXY,
YOUCOM_PROXY, BRIGHTDATA_PROXY, YOUTUBE_PROXY, INSTAGRAM_PROXY. The
instances that do not need clean egress
have no proxy: the internal searxng / crawl4ai / trafilatura (which still
take their own url/token vars) and the keyless duckduckgo.
YOUTUBE_PROXY routes the youtube reader. Where youtube.com is blocked — or
the egress IP is flagged as a
bot, which YouTube answers with "Sign in to confirm you're not a bot" — transcripts
work only through it.
INSTAGRAM_PROXY routes every request to instagram.com — the instagram
transcript reader and the instagram-profile post list alike (the latter needs
no GROQ_API_KEY). GROQ_PROXY routes the instagram reader's transcription
call to api.groq.com — an option of that instance, not a <INSTANCE>_PROXY:
Groq is its second upstream. Where Instagram is
blocked, or Groq answers Forbidden for the egress country, that leg works only
through its proxy.
The value is passed straight to httpx; socks5://host:port does proxy-side
DNS (the target hostname is resolved by the proxy, like curl --socks5-hostname), and socks5h:// / http://host:port are also accepted.
Unset → that instance goes direct. The pipeline keeps one pooled httpx client
per distinct proxy URL (and one direct client), selected per instance, so
proxied and direct providers run side by side. Needs the socks extra
(httpx[socks], already pinned).
Logging
Besides stderr (captured by Docker's rotation-capped json-file driver), the
server writes a persistent log file to data/research-mcp.log (default;
LOG_ROTATION=20 MB, LOG_RETENTION=14 days). It lives on the data/ volume,
so it survives container restarts and image updates. The file carries one
per-request line per tool call — search (query, which provider instances
actually ran, result count, latency) and read (url, the winning provider/tier
or pdf, ok, latency), plus a read_pages count=N ok=K summary — making it
useful for analyzing how requests distribute across provider tiers. No request
bodies or secrets are logged, only urls/queries, provider names, counts, timings.
Deployment
Gitea Actions builds the image and pushes it to the Gitea registry
gitea.vvzvlad.xyz/projects/research-mcp (test → build, tags latest +
sha). On prod we pull the prebuilt image via docker-compose.yml (behind
Traefik + basicAuth, watchtower auto-updates latest; the data/ volume keeps
the log file across updates) — we never build on prod.
Layout
Path | Purpose |
| Provider interfaces + |
|
|
| One module per provider type. |
| PDF detection + pypdf text extraction (used by the pipeline). |
| YouTube video-url detection + transcript fetch (the |
| Instagram url detection: post → audio transcript via Groq Whisper, profile → its posts with paging (the |
| In-code instances + pipeline order. |
| Instance loader + search/read logic (and |
|
|
| Non-secret knobs (pydantic-settings). |
|
|
| Thin entry point: build server, run streamable-http. |
| pytest suite (network mocked with respx). |
This server cannot be deployed
Maintenance
Related MCP Connectors
MCP server for 500+ pay-per-call web scraping, search, social, business, and financial data tools.
Free remote MCP server for fetching public web pages through a rotating proxy pool.
MCP server for web extraction and rendering via AceDataCloud WebExtrator
One MCP for the Web. Easily search, crawl, navigate, and extract websites without getting blocked.…
Related MCP Servers
- AlicenseAqualityCmaintenanceMCP server for web crawling, searching, and AI-powered content extraction, supporting single-page, batch, and full-site crawling along with text, news, image, book, and video search.81MIT
- AlicenseAqualityCmaintenanceMulti-source web search MCP server with RRF fusion, 4-layer URL extraction, and provider health tracking.69 npmMIT
- AlicenseAqualityDmaintenanceAn MCP server that fetches web pages and extracts clean, AI-usable context from them, enabling tools for link discovery, content search, and integrated fetch-and-search operations.58 npm1MIT
- AlicenseAqualityDmaintenanceA production-ready MCP server for free web search, news search, and webpage content extraction using DuckDuckGo and Bing.418 npmMIT