Skip to main content
Glama
competlab

competlab-mcp-server

by competlab

fetch_url

Read-only

Fetch any URL with automatic JS rendering and bot-protection handling, then return body, headers, and cleaned HTML to reduce LLM token costs.

Instructions

Fetch any URL with automatic JS-rendering and common bot-protection handling — advanced behavioral fingerprinting may still block header retrieval (surfaced via headersAvailable: false). Returns body, headers, cleanStats. Optional cleanHtml strips HTML noise while preserving text content — token-cost win for LLM consumption.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesTarget URL to fetch. Must use http:// or https:// and resolve to a public host.
cleanHtmlNoWhen `true` and the response content-type is `text/html`, strip HTML noise (scripts, styles, comments) while preserving text content. Significant token-cost reduction for LLM consumption — per-request reduction reported in `cleanStats`. Requires `bodyNeeded`.
bodyNeededNoInclude `body` and `contentType` in the response. Defaults to service-controlled value when omitted.
bodyMaxBytesNoPer-request response body cap in bytes. Accepted range 1024–104857600 (1 KiB – 100 MiB). Oversize responses are rejected pre-buffer. Defaults to service-controlled value when omitted.
maxTimeoutMsNoCaller-side timeout budget in milliseconds. Accepted range 1000–120000. Defaults to service-controlled value when omitted.
headersNeededNoInclude `headers` and `headersAvailable` in the response. Defaults to service-controlled value when omitted. When the target site uses advanced behavioral fingerprinting, `headersAvailable` is `false` and `headers` is an empty object — present, not missing. Branch on `headersAvailable`, never on whether `headers` exists.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changedv4.0.1
    • removedInput schema / additionalProperties
      Removed value: -false
  2. Changed6 schema fields changedv3.0.0
    • changedInput schema / properties / bodyMaxBytes / description
      Previous value: -"Per-request body cap in bytes. Range: 1024–104857600 (1 KiB–100 MiB)."New value: +"Per-request response body cap in bytes. Accepted range 1024–104857600 (1 KiB – 100 MiB). Oversize responses are rejected pre-buffer. Defaults to service-controlled value when omitted."
    • changedInput schema / properties / bodyNeeded / description
      Previous value: -"Include body + contentType in response. Default: true."New value: +"Include `body` and `contentType` in the response. Defaults to service-controlled value when omitted."
    • changedInput schema / properties / cleanHtml / description
      Previous value: -"Strip scripts/styles/comments from text/html responses. Requires bodyNeeded. Significant token-cost reduction for LLM consumption. Default: false."New value: +"When `true` and the response content-type is `text/html`, strip HTML noise (scripts, styles, comments) while preserving text content. Significant token-cost reduction for LLM consumption — per-request reduction reported in `cleanStats`. Requires `bodyNeeded`."
    • changedInput schema / properties / headersNeeded / description
      Previous value: -"Include headers + headersAvailable in response. Default: false. At least one of bodyNeeded or headersNeeded must be true."New value: +"Include `headers` and `headersAvailable` in the response. Defaults to service-controlled value when omitted. When the target site uses advanced behavioral fingerprinting, `headersAvailable` is `false` and `headers` is an empty object — present, not missing. Branch on `headersAvailable`, never on whether `headers` exists."
    • changedInput schema / properties / maxTimeoutMs / description
      Previous value: -"Caller timeout budget in ms. Range: 1000–120000."New value: +"Caller-side timeout budget in milliseconds. Accepted range 1000–120000. Defaults to service-controlled value when omitted."
    • changedInput schema / properties / url / description
      Previous value: -"Target URL. Must be http:// or https:// and resolve to a public host (IPv4/IPv6 literals and localhost are rejected)."New value: +"Target URL to fetch. Must use http:// or https:// and resolve to a public host."
  3. Addedv1.2.0

TDQS

A3.9/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnlyHint and openWorldHint, so the safety profile is covered. The description adds substantive behavioral context beyond that: JS-rendering, partial failure under advanced fingerprinting, and the headersAvailable:false branch signal. It stops short of covering rate limits or redirect behavior, but the failure-mode disclosure is genuinely useful.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two tightly written sentences that front-load the core capability before the caveat and the return summary. Dense but every clause carries information; no filler or repetition.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

There is no output schema, so the description carries the burden of naming return fields; it does so (body, headers, cleanStats) and explains the headersAvailable edge case. For a 6-parameter read tool this is close to complete, with only minor gaps around pagination or error shapes.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so every parameter is already documented in the schema, including the cleanHtml/bodyNeeded dependency and the headersAvailable branching rule. The description restates the cleanHtml token-cost benefit but adds no syntax or semantics the schema lacks, so the baseline of 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch any URL') and immediately layers on differentiating capabilities (automatic JS-rendering, bot-protection handling) that go beyond a tautological restatement of the name. No sibling tool performs URL fetching, so no confusion risk exists and an agent can select it unambiguously.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is implied by the description (fetching arbitrary web content for LLM consumption), but there is no explicit when-to-use/when-not-to-use statement and no named alternative, even though check_sitemap and check_ai_crawlers are network-adjacent siblings. The agent can infer the intended context but is given no routing rules.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.