Skip to main content
Glama

Clean markdown scraper

clean_markdown_scraper
Read-onlyIdempotent

Fetch a public web page by URL (or take your HTML) and return Markdown plus headings. No JavaScript rendering. Result carries a signed receipt.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlNoPublic http(s) page to fetch (ports 80/443 only). Provide exactly one of url or html.
htmlNoRaw HTML string to convert instead of fetching a URL (max 1 MiB). Provide exactly one of url or html.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
codeNo
toolNo
scopeNo
validNo
reasonNo
sourceNo
headingsNo
markdownYes
retryableNo
wordCountYes
fetched_onNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed12 schema fields changed
    • changedInput schema / examples
      Previous value: -[
      -  {
      -    "html": "<h1>Protocol Overview</h1><p>Agentic non-custodial settlement gateway.</p><ul><li>Zero custody</li><li>Sub-cent fees</li></ul>"
      -  }
      -]New value: +[
      +  {
      +    "url": "https://example.com/"
      +  }
      +]
    • changedInput schema / properties / html / description
      Previous value: -"Raw HTML string to scrape and clean into pristine Markdown."New value: +"Raw HTML string to convert instead of fetching a URL (max 1 MiB). Provide exactly one of url or html."
    • addedInput schema / properties / url
      Added value: +{
      +  "description": "Public http(s) page to fetch (ports 80/443 only). Provide exactly one of url or html.",
      +  "type": "string"
      +}
    • changedInput schema / required
      Previous value: -[
      -  "html"
      -]New value: +[]
    • addedOutput schema / properties / code
      Added value: +{
      +  "type": "string"
      +}
    • addedOutput schema / properties / fetched_on
      Added value: +{
      +  "type": "string"
      +}
    • addedOutput schema / properties / reason
      Added value: +{
      +  "type": "string"
      +}
    • addedOutput schema / properties / retryable
      Added value: +{
      +  "type": "boolean"
      +}
    • addedOutput schema / properties / scope
      Added value: +{
      +  "type": "string"
      +}
    • addedOutput schema / properties / source
      Added value: +{
      +  "type": "object"
      +}
    • addedOutput schema / properties / tool
      Added value: +{
      +  "type": "string"
      +}
    • addedOutput schema / properties / valid
      Added value: +{
      +  "type": "boolean"
      +}
  2. First observed

TDQS

A4.2/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already cover read-only, open-world and idempotent, so the bar is lower. The description still adds real behavioral context beyond them: the no-JS constraint, the fact that output includes headings, and that the result carries a signed receipt (which pairs with the receipt_verify sibling). Rate limits and failure modes are not disclosed.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Three short sentences, front-loaded with the core action, followed by the key limitation and one notable output trait. Every sentence carries information; nothing is redundant with the title.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness4/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists, so return values need not be spelled out, and the description still flags the two most relevant payload traits (Markdown + headings, signed receipt). The only gap is that it does not say what happens when both url and html are omitted, given that neither is required.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100% and both parameter descriptions already state the 'provide exactly one of url or html' rule plus the 1 MiB / port constraints. The description only restates the url-or-html duality, adding no syntax or format detail beyond the schema, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

Specific verb and resource ('Fetch a public web page by URL ... return Markdown plus headings'), with the alternate input mode named explicitly. No sibling tool in the list overlaps with web-to-Markdown conversion, so no differentiation is needed and none is missing.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines4/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

'No JavaScript rendering' is an effective negative guideline: an agent can immediately tell this tool is unsuitable for JS-rendered pages. It also names the two input modes, but stops short of naming an alternative for the JS case or stating exclusions like auth-required pages.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.