Skip to main content
Glama

fetch_page

Retrieve public web pages as clean markdown or text with metadata, escalating from HTTP to a browser only when needed and returning clear failure reasons.

Instructions

Fetch a public web page and return clean content plus metadata.

Escalation: site adapter API -> HTTP with realistic headers -> client/UA rotation -> AMP/RSS alternates -> headless browser (if installed). Every rung is listed in fetch_attempts. On failure ok=false and blocked_reason explains why (e.g. cloudflare_challenge, captcha, verification_wall, login_required, paywall, rate_limited, not_found, soft_404, js_required).

Args: url: page URL (http/https). format: which content field(s) to return. use_browser: auto (only when cheaper rungs fail), never, or always. adapter: auto, none, wechat, github, x. max_chars: truncate content to this many characters (0 = no limit). min_chars: extracted length that counts as success.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYes
formatNomarkdown
adapterNoauto
max_charsNo
min_charsNo
use_browserNoauto

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.8/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full burden and does so well: it exposes fetch_attempts, the ok=false failure contract, an enumerated blocked_reason taxonomy, and the precise semantics of use_browser 'auto'. It omits rate-limit/retry timing and permission expectations beyond the implicit 'public' scope, keeping it short of a 5.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded one-line purpose followed by a compact escalation ladder and a scannable arg list; each block earns its place. The arrow-chain is slightly dense but conveys the retry order economically.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter tool with an output schema, the description covers the behavioral contract and all inputs; return shape is correctly left to the output schema. Remaining gaps are minor (no mention of content-length limits beyond max_chars or how fetch_attempts is surfaced in the response).

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0%, so the description must compensate and largely does: it documents all six parameters, adds meaning for min_chars ('length that counts as success'), max_chars ('0 = no limit'), and lists concrete adapter values the schema itself does not enumerate. url and format are fairly thin, so not a full 5.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('Fetch a public web page') and names the return payload ('clean content plus metadata'). The 'public' qualifier implicitly brackets the tool, though it never names the sibling extract_from_html to draw a clean boundary.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The escalation ladder explains the tool's internal strategy well, which implies 'just call it and let it try everything', but there is no explicit when-to-use versus extract_from_html or when-not-to-use (e.g. authenticated/private pages). Usage context is implied rather than stated.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools