extract_page
Fetch structured web content for LLMs: extracts title, description, headings, links, images, and body as clean Markdown from any URL.
Instructions
Extract structured content from a web page: title, description, headings, links, images and the page body as clean Markdown — ready to feed to an LLM. This READS the page (use render_screenshot to SEE it).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | No | The page to read, e.g. https://example.com/blog/post | |
| html | No | Raw HTML to extract from instead of a URL | |
| format | No | json (default): full structured data; markdown: just the page body as Markdown |