fetch_html_to_text
Fetches a webpage and converts HTML to clean plain text, retaining block structure as newlines and stripping scripts, styles, and comments.
Instructions
GET a URL, decode the HTML, and return plain text with block-level structure preserved as newlines. Scripts, styles, and comments stripped; HTML entities decoded. Lighter than markdown when you only need the reading content.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL to fetch | |
| max_bytes | No | Max response size in bytes (default 5MiB) | |
| timeout_ms | No | Timeout in ms for the request, covering DNS, every redirect hop and the body (default 10000) | |
| user_agent | No | User-Agent override | |
| max_redirects | No | Max redirect hops (default 5) | |
| allow_private_hosts | No | Allow loopback / private / link-local targets for this call (default false). Refused unless the server operator launched fetch-mcp with FETCH_MCP_ALLOW_PRIVATE_HOSTS=1 -- SSRF protection stays on by default either way. |