read_url
Extract the main content from any URL as clean markdown, removing boilerplate and handling JavaScript-rendered pages. Filter by query and cache results for fast re-reads.
Instructions
Read one web page; return its main content as clean, budgeted markdown.
JS-rendered pages are handled by a real headless browser. Boilerplate
(nav, footer, ads) is stripped; if query is given, only passages relevant
to it are returned. Pages are cached — re-reads with a different query are
instant and cost no network.
Args:
url: absolute http(s) URL
query: optional focus; return only passages relevant to it
max_chars: output character budget (300-50000)
refresh: ignore cache and re-fetch the page
find: search the cached RAW text for this exact substring (case-insensitive):
returns matches with counters and context, no network needed. Requires
the page to have been read before; combine with query for first reads.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| find | No | ||
| query | No | ||
| refresh | No | ||
| max_chars | No |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| result | Yes |