io.github.VelvetSP/web-retrieval-mcp
Server Configuration
Describes the environment variables required to run the server.
| Name | Required | Description | Default |
|---|---|---|---|
| PUBMED_EMAIL | Yes | Your email address (required by NCBI) | |
| PUBMED_API_KEY | No | Optional API key for higher rate limits |
Instructions
Guidance the server publishes about itself, which clients place ahead of the tool catalog so the model reads it before choosing anything.
This server publishes no instructions, or was last inspected before Glama recorded them.
Capabilities
Features and capabilities supported by this server
Protocol revision2025-11-25
| Capability | Details |
|---|---|
| tools | {
"listChanged": false
} |
| prompts | {
"listChanged": false
} |
| resources | {
"subscribe": false,
"listChanged": false
} |
| experimental | {} |
Tools
Functions exposed to the LLM to take actions
| Name | Description |
|---|---|
| web_searchA | Search the web via Exa or Tavily. Returns one block per result, each with its OWN title, URL, highlights, and text — never a merged summary — plus a Sources trailer. In clients with built-in web search, use this as an independent complementary retrieval lane; when built-in search is disabled or unavailable, use it as the primary search path. Args:
query: the search query.
num_results: how many results (default 8).
mode: search mode — "auto" (default), "fast" (~450ms), "instant" (~250ms),
or the deep-research family "deep-lite" (≈4s), "deep" ($12/1k), or
"deep-reasoning" (12–40s, $15/1k). Deep modes request a synthesized
answer via Exa's |
| web_fetchA | Fetch a single URL's readable content. Full-body tier chain:
Exa contents → local Camoufox → optional Tavily Extract → Firecrawl.
Firecrawl is the paid last resort:
automatic full-body retrieval reaches it only after the local browser fails or
returns a body shorter than the useful-content floor. Camoufox is the one tier
that runs a real browser locally. Caller-supplied private URLs are refused before
tier selection in every render mode.
Eligible successful auto/never calls also use a host-wide 24-hour completed-result
cache in a private local Valkey sidecar. Local replay is disclosed separately from
the original provider's cache state; pass max_age_hours=0 to force provider work.
Signed/credential URLs, userinfo, positive freshness, max_age_hours=0, and
render="always" bypass completed replay. Every request is SSRF-validated before
cache access, and a hit is validated again immediately before its body is returned.
Returns content with a Args:
url: the URL to fetch.
render: "auto" (default) → Exa, then local Camoufox, optional Tavily,
then Firecrawl. "never" skips Camoufox. "always" forces Camoufox
first and skips Exa; Tavily and Firecrawl remain backstops.
For mode="concise" or a question, auto intentionally uses Exa then
Firecrawl without launching Camoufox: the browser returns a full body
and cannot satisfy the promised summary/direct-answer shape.
max_chars: max characters to request/return. None (default) → 20000-char
budget, Exa-first. An explicit value ≤10000 stays Exa-first. For an
explicit value >10000 in full-body auto, Exa cannot meet the requested
size, so the order is Camoufox → optional Tavily → Firecrawl → Exa as a
final transparently truncated salvage tier. Browser-free full-body mode
starts with optional Tavily, then Firecrawl and Exa. Semantic requests remain Exa-first
because they use Exa's summary API rather than its capped body output.
Clamped to 1000–100000. When output is (or may be) clipped, a
SSRF note: Camoufox follows redirects and re-resolves DNS, so every Camoufox attempt (automatic or render="always") is guarded:
|
| research_papersA | Search 3M+ arXiv AI/ML papers via the Firecrawl Research Index (state-of-the-art paper recall — far better than general web search for finding the right literature). Returns ranked papers: title, arXiv id, relevance score, abstract. Then call research_paper(paper_id, query=…) to verify a claim against full text before citing. SCOPE: arXiv-scoped, i.e. effectively AI/ML. For scholarly literature outside that scope (medicine, law, economics, humanities), use web_search(category="publication") — Exa's publications index (~350M works) covers what this one cannot. Args: query: natural-language research query (topic, method, benchmark, author). k: number of papers (1–25, default 8). |
| research_paperA | Inspect ONE paper from the Research Index by id (a paperId, or an arXiv id like "arxiv:2606.01509"). Without query → metadata (title, authors, categories, dates, abstract). With query → ALSO the top full-text passages answering it — use this to VERIFY a paper actually contains a method/dataset/result before relying on it. Args: paper_id: paperId or primaryId ("arxiv:NNNN.NNNNN") from research_papers. query: optional question; when set, returns claim-verification passages. |
| research_similarA | Expand from a seed paper to related work via the Research Index. SCOPE: arXiv-scoped like research_papers — for non-AI/ML literature use web_search(category="publication"). Args: paper_id: paperId or "arxiv:…" of the seed paper. intent: natural-language description of the kind of related work wanted. k: number of related papers (1–25, default 8; API allows up to 500). mode: "similar" (default), "citers" (papers citing this one), or "references" (papers this one cites). Unknown → "similar". min_score: renderer-side relevance floor (default 0.0 = off). k=8 already trims the low-score tail; raise this to filter more aggressively. rerank: optional bool; omitted from the request when None (API default is undocumented). Set True/False to force. (anchor — repeatable seed expansion — is not exposed; future work.) |
| research_githubA | Search developer primary sources via the Firecrawl Developer Index: GitHub issues, merged pull requests and repository READMEs, PLUS curated documentation sites. Returns the matched passages in Markdown, so tables and code blocks survive. Use to find the CODE behind a paper, the issue where a bug was reported and fixed, an API contract, or the discussion behind an error message. Args:
query: natural-language query (method, kernel, repo topic, error message).
k: number of results (1–25, default 8).
passages: matched passages per result (1–5, default 2).
types: restrict to any of exactly "doc", "issue", "pull_request", "readme".
NOTE the request spelling is Falls back to the de-documented legacy |
Prompts
Interactive templates invoked by user choice
| Name | Description |
|---|---|
No prompts | |
Resources
Contextual data attached and managed by the client
| Name | Description |
|---|---|
No resources | |
TDQS
Scored across 6 tools
Each tool targets a distinct retrieval surface: general web search, single-page fetch, arXiv paper search, single-paper inspection, related-paper expansion, and developer-index search. Potential overlaps are explicitly separated by corpus and use case, so an agent can reliably pick the right tool.
All names are lowercase snake_case and grouped by domain (web_* and research_*), which makes them predictable. The minor inconsistency is that web_search/web_fetch are verb-object names, while the research_* tools are object-oriented and do not encode an action as clearly.
Six tools is well-scoped for a retrieval server: two general web operations plus a four-tool research cluster. Each tool has a distinct responsibility, and there are no redundant or filler tools.
The surface covers general web search, full-page fetching, scholarly paper search, single-paper verification, related-paper exploration, and developer-source search. These cover the main retrieval workflows one would expect, with no obvious dead ends or missing core operations.