clean_web_content
Scrapes and converts any target web page into clean, LLM-ready structured Markdown, stripping ads, cookie banners, navigation clutter, modals, and script noise.
Instructions
Scrapes and converts any target web page into clean, LLM-ready structured Markdown, stripping ads, cookie banners, navigation clutter, modals, and script noise.
Usage Guidelines:
Use this tool to ingest real-time web articles, blogs, and documentation into LLM context windows.
Returns: Clean markdown body, page title, word count, and extraction metadata.
Do NOT use for YouTube video parsing (use
clean_youtube_transcript).Do NOT use for PDF whitepapers or academic papers (use
clean_pdf_research).Do NOT use for paywalled, login-required, or bot-blocked sites.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The target HTTP or HTTPS website URL to scrape and convert to markdown. | |
| auth_token_or_tx | No | Optional x402 micropayment authorization token or EVM transaction hash. | |
| respect_robots_txt | No | Whether to enforce target domain robots.txt Disallow rules (compliance-mode). |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |