Skip to main content
Glama

read_web_page

Read-onlyIdempotent

Extract clean, readable Markdown and metadata from any public webpage for LLM ingestion, stripping ads, popups, and navigational clutter. Returns clean markdown, title, description, character count, and estimated tokens. Requires x402 micropayment (0.005 USDC on Base).

When to use: Ingesting articles, blog posts, documentation, or news pages into LLM context. When NOT to use: Do NOT use for raw binary files (PDF/images), authenticated pages behind a login, or single-page apps that require heavy JavaScript rendering.

Parameters:

  • url (string, required): Full target webpage URL (e.g. 'https://news.ycombinator.com').

  • include_links (boolean, optional, default true): Whether to preserve markdown hyperlinks.

  • include_images (boolean, optional, default false): Whether to preserve image markdown links.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesFull target webpage URL to extract markdown from.
include_linksNoPreserve markdown hyperlinks.
include_imagesNoPreserve markdown image tags.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesTarget webpage URL
titleNoExtracted page title
lengthYesContent length in characters
contentYesClean extracted Markdown content
elapsed_msNoExtraction time in milliseconds
descriptionNoExtracted meta description
estimated_tokensNoEstimated LLM tokens (~len/4)

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • addedInput schema / additionalProperties
      Added value: +false
  2. Added

TDQS

A4.1/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly and idempotent, but the description adds significant behavioral detail: it strips ads/popups/navigation, returns specific fields (markdown, title, description, character count, estimated tokens), and reveals a critical x402 micropayment requirement. This goes well beyond what annotations and schema provide.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness4/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is well-organized with front-loaded purpose, return info, and payment notice, followed by clear usage guidance. The parameter bullet list is somewhat redundant with the schema, which prevents a perfect score, but overall it remains efficient and scannable.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a read-only, three-parameter tool with an output schema, the description covers the essential context: behavior, return content, payment cost, use cases, and exclusions. There are no critical gaps that would prevent an agent from calling the tool correctly.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Input schema coverage is 100%, so the schema fully documents all three parameters. The description's parameter list largely repeats schema descriptions and adds no new meaning, such as URL format constraints or interaction between include_links and include_images.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: extract clean Markdown and metadata from public webpages, with a clear intended use of LLM ingestion. It clearly conveys the tool's scope but does not explicitly name or differentiate sibling tools such as fetch_stealth_web or extract_json_from_web.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit 'When to use' and 'When NOT to use' guidance, covering article/blog ingestion and excluding binary files, authenticated pages, and JS-heavy SPAs. It lacks named alternatives, so it doesn't fully route the agent to a specific sibling tool, but the context is clear enough.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources