extract_web
Fetch clean, readable text from any public web page by stripping HTML, navigation, and scripts. Returns body text and HTTP status for summarization or analysis.
Instructions
Fetches and sanitizes readable text content from any public HTTP or HTTPS web page. Strips boilerplate HTML tags, navigation bars, and scripts. Returns clean body text and HTTP status code. Use when an agent needs primary webpage content for summarization or analysis. Do not use for authenticated pages or executing JavaScript.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The complete target website URL including http:// or https:// protocol prefix (e.g. 'https://docs.python.org/3/'). |