Skip to main content
Glama

research-mcp

A stateless MCP facade that hides a pyramid of search/read providers behind a single streamable-http MCP endpoint and exposes just 4 clean tools with good help texts. An LLM gets a simple "search → read" toolset (or both at once); behind it, several providers are tried, merged, and failed over automatically.

The app does no authentication — it is published through Traefik + basicAuth on the host. It holds no application state: the only thing persisted is a log file under data/ (kept on a volume).

Works with zero keys

make run on an empty .env is enough — no key, no SearXNG, nothing to deploy first. Two instances need no configuration at all and are therefore always enabled: duckduckgo (search — the no-JS html.duckduckgo.com SERP, one request per query, no token handshake) and trafilatura (read — local HTML→Markdown extraction). Together they are the floor: search and read both work out of the box.

Everything else is an upgrade on top of that floor. A self-hosted SearXNG (SEARXNG_URL) is optional; when it is configured it sits ahead of DuckDuckGo in the pipeline, so SearXNG keeps its copy of every url both of them return and DuckDuckGo only adds what nobody ahead of it had. (With the reranker on — the default once JINA_API_KEY is set — the merged list is reordered by relevance across all sources before the trim, so the final ordering can still change.) Every paid vendor lights up the moment its key appears.

The floor is deliberately modest: DuckDuckGo is a scraped SERP, not an API, so it paces itself to one query per 45s (skipping, never waiting, when the slot is taken) and reports a block or a captcha as a failure rather than as an empty result set. That pace is the one measured for this upstream: DuckDuckGo blocks by IP for 7-8 minutes after a burst, and a block taken here would also silence a SearXNG that reaches DuckDuckGo over the same address.

Related MCP server: Web Search MCP

Tools

Tool

What it does

web_search(query, num_results=8, language=None)

Search across all enabled providers, merge + dedup → ranked list (title, URL, snippet). Search only.

read_page(url)

One page or PDF → clean Markdown. Auto-detects type, walks the read pipeline (light → heavy) until one succeeds.

read_pages(urls)

Up to 20 urls concurrently → {summary, pages}, where each page is {url, ok, markdown} or {url, ok, error, reason}.

search_and_read(query, num_results=5, language=None)

Search and read the top hits in one call → {summary, results}, each result a hit (title, url, snippet) plus its markdown (or error + reason). Over-fetches candidates and reads them in waves, so failed urls do not eat the quota.

The tool descriptions cross-reference each other (when to take this one, when to take another), so the model gets a routing graph instead of four independent texts. search_and_read is the default for research; web_search is for links only; read_pages / read_page are for urls that are already known.

The batch tools (read_pages, search_and_read) cap each page at READ_BATCH_MAX_CHARS and mark the cut with [содержимое обрезано на N символах]; read_page always returns the page in full.

Architecture: types + instances

Providers are plugins. We separate:

  • type — an implementation class (e.g. the searxng search provider), one per module in src/providers/, registered with @register("type").

  • instance — a configured copy of a type with its secrets/URL resolved from named environment variables (multiple instances of one type are allowed, e.g. tavily-1 / tavily-2 with different keys).

Which instances exist and the order each pipeline tries them is configured in code (src/pipeline_config.py); keys/URLs come from ENV by variable name.

  • Search pipeline (searxng → duckduckgo → brave → tavily-search → firecrawl-search → jina-search → xmlriver → parallel → octen → linkup → youcom → serper → exa): enabled instances run concurrently; results are merged and deduplicated by normalized URL (earlier pipeline position wins). Position is therefore a dedup preference, not a cost gate — every enabled instance is called on every query, so cost scales with how many keys are set. When JINA_API_KEY is set (and SEARCH_RERANK_ENABLED is not turned off), the full merged list is then reranked by jina-reranker-v3.5 so the trim to num_results keeps the most relevant hits instead of a blind pipeline-order prefix; any rerank failure falls back to the merge order. searxng, duckduckgo and brave additionally throttle themselves locally (one query per 45s, 45s and 1.1s respectively, each matching a measured upstream limit — DuckDuckGo shares SearXNG's, since both reach the same engine from the same address); when the slot is taken they skip the current search instead of waiting for it. duckduckgo is the only search instance that needs no configuration, which is why it sits directly behind searxng: it is on everywhere, but a deployment that runs SearXNG must keep SearXNG's copy of every shared url.

  • Read pipeline (youtube → instagram → instagram-profile → trafilatura → jina → crawl4ai → tavily-1 → tavily-2 → firecrawl → brightdata): here the order IS a cost gate — it stops at the first sufficient answer, and brightdata (the anti-bot unlocker) sits last so it only ever sees pages everything cheaper already bounced off. The first three are url-specific readers: each is offered only the urls it recognises, ahead of the probe, and its answer is final (see below). A single probe GET classifies the url. PDFs (Content-Type / .pdf / %PDF magic) are extracted with pypdf — except a PDF with no text layer (a scan), which falls through into the chain so the remote readers get a shot at it with their own parsers, with pypdf's notice kept as the last resort. jina's OCR tier joins that attempt only when the url path ends in .pdf and jina is keyed; for HTML, that same body is handed to trafilatura so the hot path never GETs twice, then the remaining instances are tried in order and the first to return content >= FALLBACK_MIN_CHARS wins. A YouTube video url (youtube.com/watch?v=…, /shorts/…, /live/…, /embed/…, youtu.be/…) is answered before the probe with the video's transcript — title, channel, description and the captions as timestamped paragraphs — fetched from YouTube's own player API (the unofficial ANDROID client, the same one youtube-transcript-api uses). The track in the spoken language wins — a manual one when it exists, the auto-generated one otherwise. The spoken language comes from the audio track marked "original" on an auto-dubbed video (which carries auto-generated captions for every dub), and from the single auto-generated track on a plain one. A video without captions, or a fetch that fails, falls through to the normal chain above. An Instagram video url (instagram.com/reel/…, /reels/…, /p/…, /tv/…, and /<username>/reel/…, /<username>/p/…) is answered the same way, with a transcript of its audio — author, caption and the speech as timestamped paragraphs. The post comes from one anonymous request to Instagram's web GraphQL API (the logged-out query yt-dlp uses); Groq Whisper (whisper-large-v3-turbo) fetches the post's audio-only DASH track by url and transcribes it, so the server never downloads the media itself. Enabled by GROQ_API_KEY; without it, or when the post is private, login-gated or has no video, the url falls through to the normal chain. An Instagram profile url (instagram.com/<username>/) is answered with the profile's posts, 12 at a time, from the same anonymous GraphQL API (the logged-out profile posts query): per post its date, kind (reel / video / carousel / photo), link and caption; a reel or video link can be read again for its transcript (with GROQ_API_KEY). A last line Next page: …/<username>/?after=<cursor> points to the next 12 — reading that url continues the list. Needs no key. Hashtag pages are not covered: Instagram serves them only to a logged-in account.

Cross-cutting: one transient retry (5xx / transport errors) with a short backoff; 402 (out of credits) / 429 (rate limited) are treated as a provider failure → next instance (this is what makes tavily-1 → tavily-2 fail over). Vendors that report an empty balance with some other 4xx — serper answers 400 {"message": "Not enough credits"}, octen 403 "Insufficient balance" — are recognised by the body and logged as out of credits too, so an unpaid account never reads as a broken API.

An instance is enabled only if its required env var(s) are set; otherwise it is skipped with a log line. duckduckgo and trafilatura need no config (always on); jina works keyless (its key is optional). At startup the server requires at least one search and one read instance — a condition those two always satisfy, so the check now only catches a broken pipeline_config.py.

Adding a provider

  1. Write src/providers/<type>.py with a class decorated @register("<type>") implementing SearchProvider.search(...) or ReadProvider.read(...).

  2. Import the module in src/providers/__init__.py (so the decorator runs).

  3. Add an Instance("name", "<type>", api_key_env="YOUR_ENV_NAME") line in src/pipeline_config.py and reference its name in SEARCH_PIPELINE / READ_PIPELINE. Use the ENV var NAME, never a value.

  4. Document the env var in .env.example.

Quick start

make install                # create .venv + install dev/test deps
cp .env.example .env        # fill in the keys you have  (shortcut: make env)
make test                   # run tests
make run                    # run the server (streamable-http on MCP_HOST:MCP_PORT, endpoint /mcp)

Configuration

All config comes from ENV / .env (see .env.example). Provider secrets/URLs are read by name in the instance loader, not declared as Settings fields. The non-secret knobs (all defaulted): MCP_HOST, MCP_PORT, LOG_LEVEL, LOG_FILE, LOG_ROTATION, LOG_RETENTION, REQUEST_TIMEOUT, FALLBACK_MIN_CHARS, READ_PAGES_CONCURRENCY, READ_BATCH_MAX_CHARS (per-page content budget of the batch tools — read_pages and search_and_read; beyond it the markdown is cut and marked, read_page is never truncated, 0 disables the cut), RETRIES, SEARCH_RERANK_ENABLED, JINA_TOKEN_BUDGET, ALLOW_PRIVATE_NETWORK (escape hatch for the SSRF guard: true lets read_page fetch private/loopback addresses, which are blocked by default). The read_pages per-call url cap is a fixed 20 (hard constant, matching the tool description) — not configurable.

Provider env vars — duckduckgo and trafilatura take none and are always on; everything below is optional on top of them: SEARXNG_URL, BRAVE_API_KEY, SERPER_API_KEY, EXA_API_KEY, JINA_API_KEY (one key enables the jina reader in keyed mode, the jina-search provider and the search reranker; the reader alone also works keyless), CRAWL4AI_URL + CRAWL4AI_TOKEN, TAVILY_1_API_KEY, TAVILY_2_API_KEY, FIRECRAWL_API_KEY. The Tavily and Firecrawl keys each enable two instances — the reader and the search provider — because both vendors sell search and extract off one key, out of one shared monthly pool. Search runs on every query and will drain that pool well before the readers do; when it runs out, both halves stop working. Keyless until registered: XMLRIVER_USER_ID + XMLRIVER_API_KEY (Yandex SERP), PARALLEL_API_KEY, OCTEN_API_KEY, LINKUP_API_KEY, YOUCOM_API_KEY, and BRIGHTDATA_API_KEY + BRIGHTDATA_ZONE. GROQ_API_KEY enables the instagram reader (Instagram transcripts); the youtube and instagram-profile readers take no key and are always on.

Proxy

Any external instance can be routed through its own SOCKS5/HTTP proxy by setting <INSTANCE>_PROXY — useful for clean egress past IP-based blocks (e.g. Cloudflare in front of Exa). Supported per instance: EXA_PROXY, BRAVE_PROXY, SERPER_PROXY, JINA_PROXY, TAVILY_1_PROXY, TAVILY_2_PROXY, FIRECRAWL_PROXY, XMLRIVER_PROXY, PARALLEL_PROXY, OCTEN_PROXY, LINKUP_PROXY, YOUCOM_PROXY, BRIGHTDATA_PROXY, YOUTUBE_PROXY, INSTAGRAM_PROXY. The instances that do not need clean egress have no proxy: the internal searxng / crawl4ai / trafilatura (which still take their own url/token vars) and the keyless duckduckgo.

YOUTUBE_PROXY routes the youtube reader. Where youtube.com is blocked — or the egress IP is flagged as a bot, which YouTube answers with "Sign in to confirm you're not a bot" — transcripts work only through it.

INSTAGRAM_PROXY routes every request to instagram.com — the instagram transcript reader and the instagram-profile post list alike (the latter needs no GROQ_API_KEY). GROQ_PROXY routes the instagram reader's transcription call to api.groq.com — an option of that instance, not a <INSTANCE>_PROXY: Groq is its second upstream. Where Instagram is blocked, or Groq answers Forbidden for the egress country, that leg works only through its proxy.

The value is passed straight to httpx; socks5://host:port does proxy-side DNS (the target hostname is resolved by the proxy, like curl --socks5-hostname), and socks5h:// / http://host:port are also accepted. Unset → that instance goes direct. The pipeline keeps one pooled httpx client per distinct proxy URL (and one direct client), selected per instance, so proxied and direct providers run side by side. Needs the socks extra (httpx[socks], already pinned).

Logging

Besides stderr (captured by Docker's rotation-capped json-file driver), the server writes a persistent log file to data/research-mcp.log (default; LOG_ROTATION=20 MB, LOG_RETENTION=14 days). It lives on the data/ volume, so it survives container restarts and image updates. The file carries one per-request line per tool call — search (query, which provider instances actually ran, result count, latency) and read (url, the winning provider/tier or pdf, ok, latency), plus a read_pages count=N ok=K summary — making it useful for analyzing how requests distribute across provider tiers. No request bodies or secrets are logged, only urls/queries, provider names, counts, timings.

Deployment

Gitea Actions builds the image and pushes it to the Gitea registry gitea.vvzvlad.xyz/projects/research-mcp (test → build, tags latest + sha). On prod we pull the prebuilt image via docker-compose.yml (behind Traefik + basicAuth, watchtower auto-updates latest; the data/ volume keeps the log file across updates) — we never build on prod.

Layout

Path

Purpose

src/providers/base.py

Provider interfaces + SearchResult / ProviderError.

src/providers/registry.py

@register decorator → REGISTRY.

src/providers/<type>.py

One module per provider type.

src/providers/pdf.py

PDF detection + pypdf text extraction (used by the pipeline).

src/providers/youtube.py

YouTube video-url detection + transcript fetch (the youtube url-specific reader).

src/providers/instagram.py

Instagram url detection: post → audio transcript via Groq Whisper, profile → its posts with paging (the instagram / instagram_profile url-specific readers).

src/pipeline_config.py

In-code instances + pipeline order.

src/pipeline.py

Instance loader + search/read logic (and search_and_read, their composition).

src/rerank.py

JinaReranker — post-merge rerank of search results.

src/settings.py

Non-secret knobs (pydantic-settings).

src/server.py

build_server() with the 4 @mcp.tool definitions.

main.py

Thin entry point: build server, run streamable-http.

tests/

pytest suite (network mocked with respx).

Maintenance

ActivityActive
ResponsivenessUnresponsive

Related MCP Connectors

Related MCP Servers