Skip to main content
Glama

extract_from_html

Extract main content and metadata from pasted or pre-loaded HTML without network requests. Parses blocks and applies multi-strategy extraction to return markdown or text.

Instructions

Extract main content + metadata from HTML you already have (no network).

Useful when the user pasted HTML, or another tool (a browser) already loaded the page. Runs the same block detection and multi-strategy extraction as fetch_page.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlNo
htmlYes
formatNomarkdown
adapterNoauto
max_charsNo

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault

No arguments

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

A3.8/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description must carry the behavioral burden. It usefully discloses the no-network constraint and that it runs the same block detection and multi-strategy extraction as fetch_page, but says nothing about truncation behavior, the adapter mechanism, or that the url parameter is only used for link resolution.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short lines, front-loaded with the core action and scoping constraint, then the usage case, then the sibling relationship. Every sentence earns its place with no redundancy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists so return values need no explanation, and the read-only nature of extraction keeps risk low. However, five parameters at 0% schema coverage means the undocumented adapter and max_chars options leave real gaps an agent cannot resolve from structured data alone.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters2/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 0% across five parameters, so the description is the only source of parameter meaning. It mentions content and metadata but explains none of url, format, adapter, or max_chars (including the 40000-char default), leaving the agent to infer behavior from terse titles and one self-evident enum.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb and resource ('extract main content + metadata from HTML') and immediately scopes it with 'you already have (no network)'. The final sentence explicitly links it to the sibling fetch_page, so an agent can tell the two apart without reading either schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Gives clear triggering conditions: the user pasted HTML, or a browser already loaded the page. The 'no network' contrast implies fetch_page is the alternative when HTML is not yet in hand, but it never states that routing condition explicitly.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Deploy Server

Other Tools