Skip to main content
Glama

extract_page

Fetch structured web content for LLMs: extracts title, description, headings, links, images, and body as clean Markdown from any URL.

Instructions

Extract structured content from a web page: title, description, headings, links, images and the page body as clean Markdown — ready to feed to an LLM. This READS the page (use render_screenshot to SEE it).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlNoThe page to read, e.g. https://example.com/blog/post
htmlNoRaw HTML to extract from instead of a URL
formatNojson (default): full structured data; markdown: just the page body as Markdown

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A4/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

It discloses the core behavioral trait that this is a read operation ('This READS the page') and not a visual render, which is meaningful in the absence of annotations. However, it does not reveal limitations such as JavaScript rendering, network/auth behavior, rate limits, or failure modes.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Two sentences with high signal density; the action and output are front-loaded, and the sibling contrast is delivered as a short parenthetical-style remark. No filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

Given no output schema, the description does a good job listing what will be returned and the two output formats, and the optional-input semantics are covered by the schema. It doesn't mention failure behaviors or JS rendering, but for a straightforward extraction tool it is nearly complete.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3. The description mostly echoes what the schema already says (html instead of URL, json vs markdown) and adds no extra semantic detail beyond the overall output list.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States the verb 'Extract', the resource 'web page', and enumerates the exact content fields (title, description, headings, links, images, body as Markdown). The closing contrast with render_screenshot ('READS' vs 'SEE') clearly differentiates it from a key sibling.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

The description implies when to use it: when you need textual or structured content 'ready to feed to an LLM'. It explicitly routes visual needs to render_screenshot, but it does not mention alternatives like ai_extract or the exact conditions for choosing html vs url beyond what the schema already says.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.