getOgMarkdown
Convert any URL's HTML into clean Markdown, stripping navigation and ads to deliver main-content prose; optionally answer a query with only relevant chunks.
Instructions
Convert any URL's HTML into clean Markdown via the OpenGraph.io API (v3 markdown endpoint). Strips navigation, ads, and boilerplate by default — the result is main-content prose, headings, links, and images ready to read or feed into another model. Use include_tags / exclude_tags to target or remove specific page sections.
LONG PAGES — prefer retrieval over truncation. Set query with chunking: true to get back only the passages that answer your question (ranked by relevance) instead of the whole page. Use max_chars to cap raw output when you genuinely need prose. chunk_size and chunk_overlap tune the split; heading_aware keeps sections intact.
EXTRAS — include_links, include_images, and include_headings return structured link, image, and outline data, which avoids a second scrape call just to enumerate them.
UNTRUSTED CONTENT — this fetches arbitrary pages. Set ai_sanitize: true when the result will be fed to a model: it scans for prompt-injection attempts and returns a safety report. ai_sanitize_mode: 'block' rejects a risky page outright (HTTP 422) rather than returning it.
The Markdown text block is capped at 6 000 characters; the full content is always available in the structured markdown field.
Pick the right tool: getOgData → Open Graph tags, social preview metadata (title, description, image, favicon) getOgMarkdown → Clean readable text / article prose — ideal for feeding into an LLM getOgScrapeData → Raw HTML — use when you need to do your own parsing or link extraction getOgExtract → Targeted elements by tag (html_elements) or named CSS selectors (selectors) getOgScreenshot → Visual capture of a page as an image getOgQuery → Natural-language question answered from page content (100–200 credits/request)
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | URL of the webpage to convert to Markdown. | |
| query | No | Natural-language question. Returns only the most relevant chunks, ranked (BM25), instead of the whole page. Requires chunking (enabled automatically when set). | |
| retry | No | Automatically retry failed requests. | |
| cache_ok | No | Use cached results. Set to false to bypass cache. Defaults to true. | |
| chunking | No | Split the Markdown into chunks. Implied by `query`. | |
| max_chars | No | Truncate the Markdown to this many characters. Prefer `query` + `chunking` when you want the relevant part of a long page rather than an arbitrary prefix. | |
| use_proxy | No | Route the request through a standard proxy. | |
| auto_proxy | No | Automatically escalate to a proxy if the direct request fails. | |
| chunk_size | No | Target characters per chunk (200–20000). Defaults to 2000. | |
| max_chunks | No | Maximum chunks to return (1–2000). Defaults to 500. | |
| accept_lang | No | Accept-Language header for the outbound request. Defaults to 'auto'. | |
| ai_sanitize | No | Scan the fetched content for prompt-injection attempts and return a safety report. | |
| full_render | No | Force full browser rendering before conversion. Rendering is applied automatically for pages detected as JavaScript-heavy; set this when that detection is insufficient. | |
| max_retries | No | Maximum number of retry attempts (1–4). Defaults to 4. | |
| query_top_k | No | How many ranked chunks to return when `query` is set. Defaults to 5. | |
| use_premium | No | Route the request through a premium proxy. | |
| exclude_tags | No | CSS selectors to remove before conversion. Supports wildcard/regex patterns. Example: ['nav', 'footer', '.sidebar', '.ad*']. | |
| include_tags | No | CSS selectors — keep only elements matching these selectors. Example: ['article', 'main', '.content'] to target the main content area only. | |
| use_superior | No | Route the request through a superior-tier proxy. | |
| chunk_overlap | No | Characters of overlap between consecutive chunks, for context. Max half of chunk_size. | |
| heading_aware | No | Split on heading boundaries where possible, so sections stay intact. Defaults to true. | |
| include_links | No | Include every hyperlink with its text and rel attributes. | |
| max_cache_age | No | Maximum cache age in milliseconds. Defaults to 432000000 (5 days). | |
| proxy_country | No | Two-letter ISO country code for geo-targeted proxy exit node. | |
| include_chunks | No | Set false to get chunk counts in `usage` without the chunk bodies. | |
| include_images | No | Include every image with its src and alt text. | |
| load_more_wait | No | Milliseconds to wait after each load_more click (0–5000). Defaults to 1500. | |
| retry_escalate | No | Escalate proxy tier on each retry attempt. Defaults to true. | |
| ai_sanitize_mode | No | 'sanitize' cleans the content, 'warn' reports without changing it, 'block' returns HTTP 422. Only takes effect when ai_sanitize is true. | |
| include_headings | No | Include the heading outline, plus table/code-block detection flags. | |
| include_markdown | No | Set false to omit the prose body — useful when you only want structure or chunks. | |
| include_metadata | No | Include page metadata (title, description, language, canonical URL). Defaults to true. | |
| load_more_clicks | No | Number of times to click the load_more_selector (1–10). Defaults to 3. | |
| load_more_scroll | No | Scroll between load_more clicks. Defaults to true. | |
| scroll_to_bottom | No | Scroll to the bottom of the page before conversion. Forces full_render. | |
| only_main_content | No | Heuristically strip navigation, header, footer, and ads, keeping only main prose content. Defaults to true server-side. Set to false to convert the full page. | |
| wait_for_selector | No | CSS selector to wait for before converting. Forces full_render. | |
| load_more_selector | No | CSS selector for a 'load more' button to click before conversion. | |
| heading_aware_level | No | Deepest heading level treated as a split boundary (1–6). Defaults to 2. | |
| load_more_item_selector | No | CSS selector for the repeating item, used to detect when clicking stopped adding content. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | ||
| debug | No | Whether rendering, a proxy, or retries were used | |
| links | No | ||
| usage | No | Character/token counts and truncation status | |
| chunks | No | Chunk objects, ranked by relevance when `query` is set | |
| images | No | ||
| length | Yes | Character count of the returned Markdown | |
| headings | No | ||
| markdown | Yes | Full Markdown content of the page | |
| metadata | No | Title, description, language, canonical and final URL | |
| ai_safety | No | Prompt-injection report; present when ai_sanitize is true | |
| request_id | No | ||
| onlyMainContent | No |