Skip to main content
Glama
nohosa001-pixel

CleanWeb x402 — Smart Web Scraping & YouTube AI Agent

clean_web_content

Scrapes and converts any target web page into clean, LLM-ready structured Markdown, stripping ads, cookie banners, navigation clutter, modals, and script noise.

Instructions

Scrapes and converts any target web page into clean, LLM-ready structured Markdown, stripping ads, cookie banners, navigation clutter, modals, and script noise.

Usage Guidelines:

  • Use this tool to ingest real-time web articles, blogs, and documentation into LLM context windows.

  • Returns: Clean markdown body, page title, word count, and extraction metadata.

  • Do NOT use for YouTube video parsing (use clean_youtube_transcript).

  • Do NOT use for PDF whitepapers or academic papers (use clean_pdf_research).

  • Do NOT use for paywalled, login-required, or bot-blocked sites.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesThe target HTTP or HTTPS website URL to scrape and convert to markdown.
auth_token_or_txNoOptional x402 micropayment authorization token or EVM transaction hash.
respect_robots_txtNoWhether to enforce target domain robots.txt Disallow rules (compliance-mode).

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
resultYes

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed5 schema fields changedv1.2.6
    • addedInput schema / properties / auth_token_or_tx / description
      Added value: +"Optional x402 micropayment authorization token or EVM transaction hash."
    • addedInput schema / properties / respect_robots_txt
      Added value: +{
      +  "default": false,
      +  "description": "Whether to enforce target domain robots.txt Disallow rules (compliance-mode).",
      +  "title": "Respect Robots Txt",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / url / description
      Added value: +"The target HTTP or HTTPS website URL to scrape and convert to markdown."
    • addedInput schema / properties / url / examples
      Added value: +[
      +  "https://en.wikipedia.org/wiki/Web_scraping",
      +  "https://news.ycombinator.com/"
      +]
    • addedInput schema / properties / url / pattern
      Added value: +"^https?:\\/\\/[^\\s/$.?#].[^\\s]*$"
  2. Addedv1.2.5

TDQS

A4.5/5.0
Behavior4/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the burden and does well: it discloses the cleaning behavior (ads, cookie banners, nav clutter stripped), the return shape, and the failure boundary (paywalled/login-required/bot-blocked sites). It stops short of mentioning that some sites may require the micropayment auth token, which the schema hint implies.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Front-loaded with the core purpose, then a tightly scoped bulleted usage block. Every line earns its place and there is no filler.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

An output schema exists and the description still summarizes the return payload (markdown body, title, word count, metadata). For a single-URL scraping tool, nothing an agent needs to invoke it correctly is missing.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so all three parameters are already documented in the schema, including the x402 token and robots.txt flag. The description adds no further parameter semantics, so baseline 3 applies.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

States a specific verb (scrapes/converts) and resource (web page), and explicitly enumerates what it strips. It distinguishes itself from siblings clean_youtube_transcript and clean_pdf_research by name.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Provides explicit when-to-use (real-time articles, blogs, docs into LLM context) and three explicit when-NOT-to-use clauses naming the correct alternative tool for each. Nothing is left to inference.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.