research-mcp
by vvzvlad
README.md
# research-mcp
A stateless **MCP facade** that hides a pyramid of search/read providers behind a
single streamable-http MCP endpoint and exposes just **4 clean tools** with good
help texts. An LLM gets a simple "search → read" toolset (or both at once);
behind it, several providers are tried, merged, and failed over automatically.
The app does **no authentication** — it is published through Traefik + basicAuth
on the host. It holds no application state: the only thing persisted is a log
file under `data/` (kept on a volume).
## Works with zero keys
`make run` on an empty `.env` is enough — **no key, no SearXNG, nothing to deploy
first.** Two instances need no configuration at all and are therefore always
enabled: **`duckduckgo`** (search — the no-JS `html.duckduckgo.com` SERP, one
request per query, no token handshake) and **`trafilatura`** (read — local
HTML→Markdown extraction). Together they are the floor: search and read both
work out of the box.
Everything else is an upgrade on top of that floor. A self-hosted **SearXNG**
(`SEARXNG_URL`) is *optional*; when it is configured it sits **ahead** of
DuckDuckGo in the pipeline, so SearXNG keeps its copy of every url both of them
return and DuckDuckGo only adds what nobody ahead of it had. (With the reranker
on — the default once `JINA_API_KEY` is set — the merged list is reordered by
relevance across all sources before the trim, so the final ordering can still
change.) Every paid vendor lights up the moment its key appears.
The floor is deliberately modest: DuckDuckGo is a scraped SERP, not an API, so it
paces itself to one query per 45s (skipping, never waiting, when the slot is
taken) and reports a block or a captcha as a failure rather than as an empty
result set. That pace is the one measured for this upstream: DuckDuckGo blocks
by IP for 7-8 minutes after a burst, and a block taken here would also silence
a SearXNG that reaches DuckDuckGo over the same address.
## Tools
| Tool | What it does |
|------|--------------|
| `web_search(query, num_results=8, language=None)` | Search across all enabled providers, merge + dedup → ranked list (title, URL, snippet). Search only. |
| `read_page(url)` | One page or PDF → clean Markdown. Auto-detects type, walks the read pipeline (light → heavy) until one succeeds. |
| `read_pages(urls)` | Up to 20 urls concurrently → `{summary, pages}`, where each page is `{url, ok, markdown}` or `{url, ok, error, reason}`. |
| `search_and_read(query, num_results=5, language=None)` | Search **and** read the top hits in one call → `{summary, results}`, each result a hit (`title`, `url`, `snippet`) plus its `markdown` (or `error` + `reason`). Over-fetches candidates and reads them in waves, so failed urls do not eat the quota. |
The tool descriptions cross-reference each other (when to take this one, when to
take another), so the model gets a routing graph instead of four independent
texts. `search_and_read` is the default for research; `web_search` is for links
only; `read_pages` / `read_page` are for urls that are already known.
The batch tools (`read_pages`, `search_and_read`) cap each page at
`READ_BATCH_MAX_CHARS` and mark the cut with `[содержимое обрезано на N
символах]`; `read_page` always returns the page in full.
## Architecture: types + instances
Providers are **plugins**. We separate:
- **type** — an implementation class (e.g. the `searxng` search provider), one
per module in `src/providers/`, registered with `@register("type")`.
- **instance** — a configured copy of a type with its secrets/URL resolved from
**named environment variables** (multiple instances of one type are allowed,
e.g. `tavily-1` / `tavily-2` with different keys).
Which instances exist and the order each pipeline tries them is configured **in
code** (`src/pipeline_config.py`); keys/URLs come **from ENV by variable name**.
- **Search pipeline** (`searxng → duckduckgo → brave → tavily-search →
firecrawl-search → jina-search → xmlriver → parallel → octen → linkup → youcom
→ serper → exa`):
enabled instances run concurrently; results are merged and deduplicated by
normalized URL (earlier pipeline position wins). Position is therefore a
**dedup preference, not a cost gate** — every enabled instance is called on
every query, so cost scales with how many keys are set. When `JINA_API_KEY` is set (and
`SEARCH_RERANK_ENABLED` is not turned off), the full merged list is then
reranked by `jina-reranker-v3.5` so the trim to `num_results` keeps the most
relevant hits instead of a blind pipeline-order prefix; any rerank failure
falls back to the merge order.
`searxng`, `duckduckgo` and `brave` additionally throttle themselves locally
(one query per 45s, 45s and 1.1s respectively, each matching a measured
upstream limit — DuckDuckGo shares SearXNG's, since both reach the same engine
from the same address); when the slot is taken they **skip** the current
search instead of waiting for it. `duckduckgo` is the only search instance
that needs no configuration, which is why it sits directly behind
`searxng`: it is on everywhere, but a deployment that runs SearXNG must
keep SearXNG's copy of every shared url.
- **Read pipeline** (`youtube → instagram → instagram-profile → trafilatura →
jina → crawl4ai → tavily-1 → tavily-2 → firecrawl → brightdata`): here the
order IS a cost gate — it stops at the first sufficient answer, and
`brightdata` (the anti-bot unlocker) sits last so it only ever sees pages
everything cheaper already bounced off. The first three are url-specific
readers: each is offered only the urls it recognises, ahead of the probe, and
its answer is final (see below). A single probe
GET classifies the url. PDFs (Content-Type /
`.pdf` / `%PDF` magic) are extracted with pypdf — except a PDF with **no text
layer** (a scan), which falls through into the chain so the remote readers get
a shot at it with their own parsers, with pypdf's notice kept as the last
resort. jina's OCR tier joins that attempt only when the url path ends in
`.pdf` *and* jina is keyed; for HTML, that same body is
handed to `trafilatura` so the hot path never GETs twice, then the remaining
instances are tried in order and the first to return content
`>= FALLBACK_MIN_CHARS` wins.
A **YouTube video url** (`youtube.com/watch?v=…`, `/shorts/…`, `/live/…`,
`/embed/…`, `youtu.be/…`) is answered before the probe with the video's
**transcript** — title, channel, description and the captions as timestamped
paragraphs — fetched from YouTube's own player API (the unofficial ANDROID
client, the same one `youtube-transcript-api` uses). The track in the spoken
language wins — a manual one when it exists, the auto-generated one otherwise.
The spoken language comes from the audio track marked "original" on an
auto-dubbed video (which carries auto-generated captions for every dub), and
from the single auto-generated track on a plain one.
A video without captions, or a fetch that fails, falls through to the normal
chain above.
An **Instagram video url** (`instagram.com/reel/…`, `/reels/…`, `/p/…`,
`/tv/…`, and `/<username>/reel/…`, `/<username>/p/…`) is answered the same way, with a
**transcript of its audio** — author, caption and the speech as timestamped
paragraphs. The post comes from one anonymous request to Instagram's web
GraphQL API (the logged-out query `yt-dlp` uses); Groq Whisper
(`whisper-large-v3-turbo`) fetches the post's audio-only DASH track by url and
transcribes it, so the server never downloads the media itself. Enabled by
`GROQ_API_KEY`; without it, or when the post is private, login-gated or has no
video, the url falls through to the normal chain.
An **Instagram profile url** (`instagram.com/<username>/`) is answered with the
profile's **posts**, 12 at a time, from the same anonymous GraphQL API (the
logged-out profile posts query): per post its date, kind (reel / video /
carousel / photo), link and caption; a reel or video link can be read again
for its transcript (with `GROQ_API_KEY`). A last line `Next page: …/<username>/?after=<cursor>` points to
the next 12 — reading that url continues the list. Needs no key. Hashtag pages
are not covered: Instagram serves them only to a logged-in account.
Cross-cutting: one transient retry (5xx / transport errors) with a short backoff;
**402 (out of credits) / 429 (rate limited) are treated as a provider failure →
next instance** (this is what makes `tavily-1 → tavily-2` fail over). Vendors
that report an empty balance with some other 4xx — serper answers `400 {"message":
"Not enough credits"}`, octen `403 "Insufficient balance"` — are recognised by
the body and logged as `out of credits` too, so an unpaid account never reads as
a broken API.
An instance is **enabled** only if its required env var(s) are set; otherwise it
is skipped with a log line. `duckduckgo` and `trafilatura` need no config (always
on); `jina` works keyless (its key is optional). At startup the server requires at
least one search and one read instance — a condition those two always satisfy, so
the check now only catches a broken `pipeline_config.py`.
## Adding a provider
1. Write `src/providers/<type>.py` with a class decorated `@register("<type>")`
implementing `SearchProvider.search(...)` or `ReadProvider.read(...)`.
2. Import the module in `src/providers/__init__.py` (so the decorator runs).
3. Add an `Instance("name", "<type>", api_key_env="YOUR_ENV_NAME")` line in
`src/pipeline_config.py` and reference its `name` in `SEARCH_PIPELINE` /
`READ_PIPELINE`. **Use the ENV var NAME, never a value.**
4. Document the env var in `.env.example`.
## Quick start
```bash
make install # create .venv + install dev/test deps
cp .env.example .env # fill in the keys you have (shortcut: make env)
make test # run tests
make run # run the server (streamable-http on MCP_HOST:MCP_PORT, endpoint /mcp)
```
## Configuration
All config comes from ENV / `.env` (see `.env.example`). Provider secrets/URLs
are read by **name** in the instance loader, not declared as Settings fields. The
non-secret knobs (all defaulted): `MCP_HOST`, `MCP_PORT`, `LOG_LEVEL`,
`LOG_FILE`, `LOG_ROTATION`, `LOG_RETENTION`, `REQUEST_TIMEOUT`,
`FALLBACK_MIN_CHARS`, `READ_PAGES_CONCURRENCY`, `READ_BATCH_MAX_CHARS` (per-page
content budget of the batch tools — `read_pages` and `search_and_read`; beyond it
the markdown is cut and marked, `read_page` is never truncated, `0` disables the
cut), `RETRIES`,
`SEARCH_RERANK_ENABLED`, `JINA_TOKEN_BUDGET`, `ALLOW_PRIVATE_NETWORK` (escape
hatch for the SSRF guard: `true` lets `read_page` fetch private/loopback
addresses, which are blocked by default). The `read_pages`
per-call url cap is a fixed `20` (hard constant, matching the tool description) —
not configurable.
Provider env vars — `duckduckgo` and `trafilatura` take none and are always on;
everything below is optional on top of them: `SEARXNG_URL`, `BRAVE_API_KEY`,
`SERPER_API_KEY`, `EXA_API_KEY`, `JINA_API_KEY`
(one key enables the `jina` reader in keyed mode, the `jina-search` provider and
the search reranker; the reader alone also works keyless), `CRAWL4AI_URL` +
`CRAWL4AI_TOKEN`, `TAVILY_1_API_KEY`, `TAVILY_2_API_KEY`, `FIRECRAWL_API_KEY`.
The Tavily and Firecrawl keys each enable **two** instances — the reader and the
search provider — because both vendors sell search and extract off one key, out
of one shared monthly pool. Search runs on every query and will drain that pool
well before the readers do; when it runs out, both halves stop working.
Keyless until registered: `XMLRIVER_USER_ID` + `XMLRIVER_API_KEY` (Yandex SERP),
`PARALLEL_API_KEY`, `OCTEN_API_KEY`, `LINKUP_API_KEY`, `YOUCOM_API_KEY`, and
`BRIGHTDATA_API_KEY` + `BRIGHTDATA_ZONE`. `GROQ_API_KEY` enables the `instagram`
reader (Instagram transcripts); the `youtube` and `instagram-profile` readers
take no key and are always on.
## Proxy
Any external instance can be routed through its own **SOCKS5/HTTP proxy** by
setting `<INSTANCE>_PROXY` — useful for clean egress past IP-based blocks (e.g.
Cloudflare in front of Exa). Supported per instance: `EXA_PROXY`, `BRAVE_PROXY`, `SERPER_PROXY`,
`JINA_PROXY`, `TAVILY_1_PROXY`, `TAVILY_2_PROXY`, `FIRECRAWL_PROXY`,
`XMLRIVER_PROXY`, `PARALLEL_PROXY`, `OCTEN_PROXY`, `LINKUP_PROXY`,
`YOUCOM_PROXY`, `BRIGHTDATA_PROXY`, `YOUTUBE_PROXY`, `INSTAGRAM_PROXY`. The
instances that do not need clean egress
have no proxy: the internal `searxng` / `crawl4ai` / `trafilatura` (which still
take their own url/token vars) and the keyless `duckduckgo`.
`YOUTUBE_PROXY` routes the `youtube` reader. Where youtube.com is blocked — or
the egress IP is flagged as a
bot, which YouTube answers with "Sign in to confirm you're not a bot" — transcripts
work only through it.
`INSTAGRAM_PROXY` routes every request to instagram.com — the `instagram`
transcript reader and the `instagram-profile` post list alike (the latter needs
no `GROQ_API_KEY`). `GROQ_PROXY` routes the `instagram` reader's transcription
call to api.groq.com — an option of that instance, not a `<INSTANCE>_PROXY`:
Groq is its second upstream. Where Instagram is
blocked, or Groq answers `Forbidden` for the egress country, that leg works only
through its proxy.
The value is passed straight to httpx; `socks5://host:port` does **proxy-side
DNS** (the target hostname is resolved by the proxy, like `curl
--socks5-hostname`), and `socks5h://` / `http://host:port` are also accepted.
Unset → that instance goes direct. The pipeline keeps one pooled httpx client
per distinct proxy URL (and one direct client), selected per instance, so
proxied and direct providers run side by side. Needs the `socks` extra
(`httpx[socks]`, already pinned).
## Logging
Besides stderr (captured by Docker's rotation-capped json-file driver), the
server writes a **persistent log file** to `data/research-mcp.log` (default;
`LOG_ROTATION=20 MB`, `LOG_RETENTION=14 days`). It lives on the `data/` volume,
so it survives container restarts and image updates. The file carries one
**per-request line** per tool call — search (`query`, which provider instances
actually ran, result count, latency) and read (`url`, the winning provider/tier
or `pdf`, `ok`, latency), plus a `read_pages count=N ok=K` summary — making it
useful for analyzing how requests distribute across provider tiers. No request
bodies or secrets are logged, only urls/queries, provider names, counts, timings.
## Deployment
Gitea Actions builds the image and pushes it to the Gitea registry
`gitea.vvzvlad.xyz/projects/research-mcp` (`test` → `build`, tags `latest` +
`sha`). On prod we pull the prebuilt image via `docker-compose.yml` (behind
Traefik + basicAuth, watchtower auto-updates `latest`; the `data/` volume keeps
the log file across updates) — we never build on prod.
## Layout
| Path | Purpose |
|------|---------|
| `src/providers/base.py` | Provider interfaces + `SearchResult` / `ProviderError`. |
| `src/providers/registry.py` | `@register` decorator → `REGISTRY`. |
| `src/providers/<type>.py` | One module per provider type. |
| `src/providers/pdf.py` | PDF detection + pypdf text extraction (used by the pipeline). |
| `src/providers/youtube.py` | YouTube video-url detection + transcript fetch (the `youtube` url-specific reader). |
| `src/providers/instagram.py` | Instagram url detection: post → audio transcript via Groq Whisper, profile → its posts with paging (the `instagram` / `instagram_profile` url-specific readers). |
| `src/pipeline_config.py` | In-code instances + pipeline order. |
| `src/pipeline.py` | Instance loader + search/read logic (and `search_and_read`, their composition). |
| `src/rerank.py` | `JinaReranker` — post-merge rerank of search results. |
| `src/settings.py` | Non-secret knobs (pydantic-settings). |
| `src/server.py` | `build_server()` with the 4 `@mcp.tool` definitions. |
| `main.py` | Thin entry point: build server, run streamable-http. |
| `tests/` | pytest suite (network mocked with respx). |
This server cannot be deployed
Maintenance
ActivityActive
ResponsivenessUnresponsive