Fetch a URL as Markdown
fetch_urlFetch a web page or document and convert it to clean Markdown by removing navigation, ads, and cookie banners; supports HTML, PDF, Office, EPUB, CSV, JSON, and text.
Instructions
Fetch one web page or document and return its main content as clean Markdown, with the navigation, adverts and cookie banners removed. Handles HTML (via a readability pass), PDF, DOCX, XLSX, PPTX, EPUB, ODT, CSV, JSON and plain text. Results are stored in a local SQLite index, so reading the same URL again is instant and the page becomes searchable offline with search_index. Long pages are truncated rather than refused: the response reports a next offset you can pass back to continue. robots.txt is honoured by default and private/loopback addresses are refused as a safety measure.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Absolute http(s) URL to fetch. | |
| format | No | Output format. "markdown" (default) is compact and readable; "json" returns structured data for post-processing. | |
| offset | No | Skip this many characters into the content. Use the `next offset` from a previous call to page through a long document. | |
| refresh | No | Ignore the local cache and re-fetch (default false). | |
| max_chars | No | Maximum characters of content to return. Longer content is truncated and a `next offset` is reported so you can continue reading. | |
| use_cache | No | Use the cached copy when it is fresh (default true). | |
| include_links | No | Also return the outgoing links found in the body, for follow-up fetching. | |
| respect_robots | No | Honour robots.txt for this request (default: the server setting, which is true). |