Skip to main content
Glama
VelvetSP

io.github.VelvetSP/web-retrieval-mcp

by VelvetSP

web_fetch

Read-only

Fetch readable content from a given URL using a tiered fallback of retrieval engines, with caching and SSRF protection for reliable, secure access.

Instructions

Fetch a single URL's readable content. Full-body tier chain: Exa contents → local Camoufox → optional Tavily Extract → Firecrawl. Firecrawl is the paid last resort: automatic full-body retrieval reaches it only after the local browser fails or returns a body shorter than the useful-content floor. Camoufox is the one tier that runs a real browser locally. Caller-supplied private URLs are refused before tier selection in every render mode. Eligible successful auto/never calls also use a host-wide 24-hour completed-result cache in a private local Valkey sidecar. Local replay is disclosed separately from the original provider's cache state; pass max_age_hours=0 to force provider work. Signed/credential URLs, userinfo, positive freshness, max_age_hours=0, and render="always" bypass completed replay. Every request is SSRF-validated before cache access, and a hit is validated again immediately before its body is returned. Returns content with a [served by: …] provenance header. In clients with built-in page retrieval, use this as an independent complementary retrieval lane; when built-in retrieval is disabled or unavailable, use it as the primary fetch path.

Args: url: the URL to fetch. render: "auto" (default) → Exa, then local Camoufox, optional Tavily, then Firecrawl. "never" skips Camoufox. "always" forces Camoufox first and skips Exa; Tavily and Firecrawl remain backstops. For mode="concise" or a question, auto intentionally uses Exa then Firecrawl without launching Camoufox: the browser returns a full body and cannot satisfy the promised summary/direct-answer shape. max_chars: max characters to request/return. None (default) → 20000-char budget, Exa-first. An explicit value ≤10000 stays Exa-first. For an explicit value >10000 in full-body auto, Exa cannot meet the requested size, so the order is Camoufox → optional Tavily → Firecrawl → Exa as a final transparently truncated salvage tier. Browser-free full-body mode starts with optional Tavily, then Firecrawl and Exa. Semantic requests remain Exa-first because they use Exa's summary API rather than its capped body output. Clamped to 1000–100000. When output is (or may be) clipped, a [TRUNCATED at N chars — …] marker is appended on its own line. max_age_hours: freshness window for the Exa AND Firecrawl tiers. None = each tier's default cache (Exa default; Firecrawl ~2 days); 0 = force fresh on both. -1 = "always use cache" (Exa-documented) — Firecrawl has NO equivalent, so on that tier -1 is treated as omit (its own default cache window applies instead; disclosed in the [cache: …] line). Values below -1 are ignored (stderr note, default cache used). Values above 720 (Exa's documented ceiling) are CLAMPED to 720, and the clamp is disclosed — a clamped value changes which content can come back. The camoufox render tier is always live. When Firecrawl serves with cache permitted, a [cache: …] disclosure line is added. mode: "full" (default, whole readable body) or "concise" (a generated summary — far fewer tokens). Honored by the Exa and Firecrawl tiers; the camoufox render tier ignores it (returns full body, no [mode:] line). Concise/question outputs carry a [mode: …] provenance line. question: optional grounded-extraction query. When set, the tier returns a direct ANSWER to the question (Exa summary-with-query / Firecrawl question format) instead of the page body; short answers are accepted — the floor is 1 char, so only an EMPTY answer cascades to the next tier ("Paris"/"No" are legitimate answers). The 60-char extract floor applies to mode="concise", and the 200-char floor to a full body. Overrides mode. tavily: enable or disable Tavily Extract for this call. When omitted, WEB_FETCH_TAVILY_TIER is parsed strictly (default false). Tavily is a full-body tier only and is skipped for concise/question requests and whenever max_age_hours is explicit because it cannot honor those contracts. Requires the tavily extra and TAVILY_API_KEY.

SSRF note: Camoufox follows redirects and re-resolves DNS, so every Camoufox attempt (automatic or render="always") is guarded:

  • _make_route_guard aborts any request whose host resolves non-public, classifying by RESOLVED IP (not the URL string), failing closed on a resolution error;

  • _make_request_observer + _flush_pending exist because page.route does NOT fire on a main-frame 3xx — they see the redirect hops the route guard alone would miss;

  • _camoufox_render raises on a blocked hop at FOUR checkpoints: after a goto exception (catches a navigation the guard itself aborted, raising the guard's reason instead of an opaque playwright error), after a successful goto, after the networkidle wait, and after inner_text (a hop recorded during the extraction await, before any body returns).

  • test_ssrf_redirect_live.py covers this live. The observer detects a forbidden document request after Chromium has emitted it, so the application prevents private content from being returned but cannot prove that no outbound packet was sent. Chromium may also re-resolve after the guard's check. Full closure would need a validating forward proxy or equivalent network policy.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYes
modeNofull
renderNoauto
tavilyNo
questionNo
max_charsNo
max_age_hoursNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.2.0

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotatations only provide readOnlyHint, but the description discloses the full tier chain, cache behavior, bypass conditions, SSRF validation, truncation markers, and provenance headers. It even goes as far as admitting the residual limit that an outbound packet cannot be fully ruled out, which is exceptional transparency beyond the annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long but well-structured and front-loaded: a one-sentence purpose, a tier-chain overview, then an Args section with each parameter explained. The density is justified by the tool's complexity, and the organization allows an agent to extract common-case usage quickly while still having edge-case details available.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given 7 parameters, 0% schema coverage, and safety-sensitive SSRF behavior, the description fully covers purpose, parameter semantics, caching, render modes, return provenance, and security limitations. No essential decision an agent needs to make before invoking the tool is left undocumented.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description carries the full burden. It thoroughly documents all seven parameters, including defaults, clamped ranges, tier-order changes, cross-parameter interactions, and edge cases like max_age_hours=-1 and the >10000 max_chars behavior.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description opens with 'Fetch a single URL's readable content'—a specific verb, resource, and output type. It clearly distinguishes web_fetch from the research_* and web_search siblings by scoping it to one URL and describing the readable-content retrieval role.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Explicitly states when to use this tool as a complementary retrieval lane versus the primary fetch path, and details render-mode selection rules (auto/never/always). It also explains when individual tiers are skipped, such as Tavily being omitted for concise/question requests, which gives clear when-to-use and when-not-to-use guidance.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.