Skip to main content
Glama
VelvetSP

io.github.VelvetSP/web-retrieval-mcp

by VelvetSP

web_search

Read-only

Search the web using Exa or Tavily, returning individual result blocks with titles, URLs, and highlights. Supports filters like recency, category, and domain inclusion/exclusion.

Instructions

Search the web via Exa or Tavily. Returns one block per result, each with its OWN title, URL, highlights, and text — never a merged summary — plus a Sources trailer. In clients with built-in web search, use this as an independent complementary retrieval lane; when built-in search is disabled or unavailable, use it as the primary search path.

Args: query: the search query. num_results: how many results (default 8). mode: search mode — "auto" (default), "fast" (~450ms), "instant" (~250ms), or the deep-research family "deep-lite" (≈4s), "deep" ($12/1k), or "deep-reasoning" (12–40s, $15/1k). Deep modes request a synthesized answer via Exa's outputSchema and return a "## Synthesized answer" block — with a "Grounding:" citation line when Exa returns one — plus per-result blocks; this costs ~2s of synthesis latency on top of the mode's own search time, which is why it is scoped to the deep family only (their timeout budget already covers it). Legacy "neural"/"keyword" map to "auto" (deprecated). text_chars: per-result body length (default 1200; clamped 200–8000). Now governs BOTH the Exa request size (request-what-you-render) and the rendered cap — raise it to surface more body text per result. recency_days: only results published within the last N days (maps to startPublishedDate). Pass this for time-sensitive/"latest" queries — the tool does NOT auto-tighten dates on its own. recency_hours: like recency_days but HOUR granularity — reaches BOTH tiers (Exa via a full ISO timestamp; the Firecrawl fallback via tbs qdr:h/cdr:1,cd_min:…). Full precedence ladder, same on both tiers: explicit VALID start_published_date > recency_hours > recency_days. An invalid/absent value at each tier falls through to the next. start_published_date / end_published_date: ISO date bounds ("2026-01-15" or a full ISO timestamp). Wins over recency_hours/recency_days per the ladder above on both tiers. The fallback represents an explicit start with Firecrawl's custom-date-range syntax. Unparseable values are ignored (stderr note), never an error. category: Exa category hint ("news", "publication", "company", "people", "financial report", "personal site", or a free string). NOTE: "company" and "people" forbid date filters + excludeDomains (dropped automatically), and "people" restricts includeDomains to Exa's supported profile domains (it 400s otherwise — surfaced as SEARCH_FAILED / fallback). Enum migration: "research paper" was RENAMED to "publication" and is auto-aliased forward; "tweet" is GONE (Exa 400s) and is dropped, leaving the search unscoped; "pdf"/"github" are deprecated-but-live and pass through. Each of those emits a [category=… ] line in the response header, since dropping or renaming a category changes which results you get back. include_domains / exclude_domains: restrict/exclude result hosts (capped at 20 entries each). summary: when True, request a generated per-result summary instead of raw text/highlights (fewer tokens; text + highlights omitted from the request). sort_by_date: sort results by publish date instead of relevance. Honoured ONLY on the Firecrawl fallback tier (Firecrawl's sbd:1); Exa has no equivalent — when the Exa tier serves and this was requested, the response header discloses that it was not applied rather than silently ignoring it. provider: "exa" or "tavily". When omitted, WEB_SEARCH_PROVIDER is used, defaulting to Exa. Invalid values fail explicitly. Tavily requires the tavily extra and TAVILY_API_KEY.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
modeNoauto
queryYes
summaryNo
categoryNo
providerNo
text_charsNo
num_resultsNo
recency_daysNo
sort_by_dateNo
recency_hoursNo
exclude_domainsNo
include_domainsNo
end_published_dateNo
start_published_dateNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations only declare readOnlyHint=true, but the description discloses far more: results are never merged into a combined summary; deep modes add latency and cost; recency filters follow an explicit precedence ladder; category values like 'people' can cause 400s; and sort_by_date is silently not applied on Exa but is disclosed in the response header. These behavioral traits go well beyond the annotations and give an agent accurate expectations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but the length is warranted by 14 parameters and 0% schema description coverage. It is front-loaded with the core behavior and return format, then organized as a labeled Args list where each parameter earns its place with operational detail. A few sentences could be tightened, but none are pure filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Combined with the readOnlyHint and the presence of an output schema, the description covers return granularity, mode behavior, provider fallback, parameter interactions, failure modes, and date-filtering semantics. An agent has enough context to decide whether to call this tool and how to set the right parameters without needing to guess.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must carry the full burden, and it does. The Args section documents every parameter, including defaults, clamping ('clamped 200–8000'), precedence ('explicit VALID start_published_date > recency_hours > recency_days'), enum migrations, and provider-specific behavior. It fully compensates for the missing schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with a specific verb-resource pair — 'Search the web via Exa or Tavily' — and immediately defines the return shape: 'one block per result, each with its OWN title, URL, highlights, and text — never a merged summary'. It clearly differentiates this tool from built-in web search and from the research-paper-specific sibling tools by framing it as general-purpose web search, but it does not explicitly name sibling tools as alternatives.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description explicitly tells an agent when to use this tool: 'In clients with built-in web search, use this as an independent complementary retrieval lane; when built-in search is disabled or unavailable, use it as the primary search path.' This is a direct, actionable selection rule with respect to the most relevant alternative, and it also warns about provider requirements such as 'Tavily requires the tavily extra and TAVILY_API_KEY.'

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.