Skip to main content
Glama

getOgMarkdown

Read-onlyIdempotent

Convert any URL's HTML into clean Markdown, stripping navigation and ads to deliver main-content prose; optionally answer a query with only relevant chunks.

Instructions

Convert any URL's HTML into clean Markdown via the OpenGraph.io API (v3 markdown endpoint). Strips navigation, ads, and boilerplate by default — the result is main-content prose, headings, links, and images ready to read or feed into another model. Use include_tags / exclude_tags to target or remove specific page sections.

LONG PAGES — prefer retrieval over truncation. Set query with chunking: true to get back only the passages that answer your question (ranked by relevance) instead of the whole page. Use max_chars to cap raw output when you genuinely need prose. chunk_size and chunk_overlap tune the split; heading_aware keeps sections intact.

EXTRAS — include_links, include_images, and include_headings return structured link, image, and outline data, which avoids a second scrape call just to enumerate them.

UNTRUSTED CONTENT — this fetches arbitrary pages. Set ai_sanitize: true when the result will be fed to a model: it scans for prompt-injection attempts and returns a safety report. ai_sanitize_mode: 'block' rejects a risky page outright (HTTP 422) rather than returning it.

The Markdown text block is capped at 6 000 characters; the full content is always available in the structured markdown field.

Pick the right tool: getOgData → Open Graph tags, social preview metadata (title, description, image, favicon) getOgMarkdown → Clean readable text / article prose — ideal for feeding into an LLM getOgScrapeData → Raw HTML — use when you need to do your own parsing or link extraction getOgExtract → Targeted elements by tag (html_elements) or named CSS selectors (selectors) getOgScreenshot → Visual capture of a page as an image getOgQuery → Natural-language question answered from page content (100–200 credits/request)

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesURL of the webpage to convert to Markdown.
queryNoNatural-language question. Returns only the most relevant chunks, ranked (BM25), instead of the whole page. Requires chunking (enabled automatically when set).
retryNoAutomatically retry failed requests.
cache_okNoUse cached results. Set to false to bypass cache. Defaults to true.
chunkingNoSplit the Markdown into chunks. Implied by `query`.
max_charsNoTruncate the Markdown to this many characters. Prefer `query` + `chunking` when you want the relevant part of a long page rather than an arbitrary prefix.
use_proxyNoRoute the request through a standard proxy.
auto_proxyNoAutomatically escalate to a proxy if the direct request fails.
chunk_sizeNoTarget characters per chunk (200–20000). Defaults to 2000.
max_chunksNoMaximum chunks to return (1–2000). Defaults to 500.
accept_langNoAccept-Language header for the outbound request. Defaults to 'auto'.
ai_sanitizeNoScan the fetched content for prompt-injection attempts and return a safety report.
full_renderNoForce full browser rendering before conversion. Rendering is applied automatically for pages detected as JavaScript-heavy; set this when that detection is insufficient.
max_retriesNoMaximum number of retry attempts (1–4). Defaults to 4.
query_top_kNoHow many ranked chunks to return when `query` is set. Defaults to 5.
use_premiumNoRoute the request through a premium proxy.
exclude_tagsNoCSS selectors to remove before conversion. Supports wildcard/regex patterns. Example: ['nav', 'footer', '.sidebar', '.ad*'].
include_tagsNoCSS selectors — keep only elements matching these selectors. Example: ['article', 'main', '.content'] to target the main content area only.
use_superiorNoRoute the request through a superior-tier proxy.
chunk_overlapNoCharacters of overlap between consecutive chunks, for context. Max half of chunk_size.
heading_awareNoSplit on heading boundaries where possible, so sections stay intact. Defaults to true.
include_linksNoInclude every hyperlink with its text and rel attributes.
max_cache_ageNoMaximum cache age in milliseconds. Defaults to 432000000 (5 days).
proxy_countryNoTwo-letter ISO country code for geo-targeted proxy exit node.
include_chunksNoSet false to get chunk counts in `usage` without the chunk bodies.
include_imagesNoInclude every image with its src and alt text.
load_more_waitNoMilliseconds to wait after each load_more click (0–5000). Defaults to 1500.
retry_escalateNoEscalate proxy tier on each retry attempt. Defaults to true.
ai_sanitize_modeNo'sanitize' cleans the content, 'warn' reports without changing it, 'block' returns HTTP 422. Only takes effect when ai_sanitize is true.
include_headingsNoInclude the heading outline, plus table/code-block detection flags.
include_markdownNoSet false to omit the prose body — useful when you only want structure or chunks.
include_metadataNoInclude page metadata (title, description, language, canonical URL). Defaults to true.
load_more_clicksNoNumber of times to click the load_more_selector (1–10). Defaults to 3.
load_more_scrollNoScroll between load_more clicks. Defaults to true.
scroll_to_bottomNoScroll to the bottom of the page before conversion. Forces full_render.
only_main_contentNoHeuristically strip navigation, header, footer, and ads, keeping only main prose content. Defaults to true server-side. Set to false to convert the full page.
wait_for_selectorNoCSS selector to wait for before converting. Forces full_render.
load_more_selectorNoCSS selector for a 'load more' button to click before conversion.
heading_aware_levelNoDeepest heading level treated as a split boundary (1–6). Defaults to 2.
load_more_item_selectorNoCSS selector for the repeating item, used to detect when clicking stopped adding content.

Output Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYes
debugNoWhether rendering, a proxy, or retries were used
linksNo
usageNoCharacter/token counts and truncation status
chunksNoChunk objects, ranked by relevance when `query` is set
imagesNo
lengthYesCharacter count of the returned Markdown
headingsNo
markdownYesFull Markdown content of the page
metadataNoTitle, description, language, canonical and final URL
ai_safetyNoPrompt-injection report; present when ai_sanitize is true
request_idNo
onlyMainContentNo

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed31 schema fields changedv2.1.0
    • changedInput schema / properties / ai_sanitize / description
      Previous value: -"Scan the fetched content for prompt-injection attempts."New value: +"Scan the fetched content for prompt-injection attempts and return a safety report."
    • changedInput schema / properties / ai_sanitize_mode / description
      Previous value: -"'sanitize' cleans the content, 'warn' returns a safety report, 'block' returns HTTP 422."New value: +"'sanitize' cleans the content, 'warn' reports without changing it, 'block' returns HTTP 422. Only takes effect when ai_sanitize is true."
    • addedInput schema / properties / chunk_overlap
      Added value: +{
      +  "description": "Characters of overlap between consecutive chunks, for context. Max half of chunk_size.",
      +  "minimum": 0,
      +  "type": "integer"
      +}
    • addedInput schema / properties / chunk_size
      Added value: +{
      +  "description": "Target characters per chunk (200–20000). Defaults to 2000.",
      +  "maximum": 20000,
      +  "minimum": 200,
      +  "type": "integer"
      +}
    • addedInput schema / properties / chunking
      Added value: +{
      +  "description": "Split the Markdown into chunks. Implied by `query`.",
      +  "type": "boolean"
      +}
    • changedInput schema / properties / full_render / description
      Previous value: -"Fully render the page with JavaScript before conversion. REQUIRED for SPAs and JS-heavy sites — v3 auto_render does NOT apply to the markdown pipeline."New value: +"Force full browser rendering before conversion. Rendering is applied automatically for pages detected as JavaScript-heavy; set this when that detection is insufficient."
    • addedInput schema / properties / heading_aware
      Added value: +{
      +  "description": "Split on heading boundaries where possible, so sections stay intact. Defaults to true.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / heading_aware_level
      Added value: +{
      +  "description": "Deepest heading level treated as a split boundary (1–6). Defaults to 2.",
      +  "maximum": 6,
      +  "minimum": 1,
      +  "type": "integer"
      +}
    • addedInput schema / properties / include_chunks
      Added value: +{
      +  "description": "Set false to get chunk counts in `usage` without the chunk bodies.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / include_headings
      Added value: +{
      +  "description": "Include the heading outline, plus table/code-block detection flags.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / include_images
      Added value: +{
      +  "description": "Include every image with its src and alt text.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / include_links
      Added value: +{
      +  "description": "Include every hyperlink with its text and rel attributes.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / include_markdown
      Added value: +{
      +  "description": "Set false to omit the prose body — useful when you only want structure or chunks.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / include_metadata
      Added value: +{
      +  "description": "Include page metadata (title, description, language, canonical URL). Defaults to true.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / load_more_item_selector
      Added value: +{
      +  "description": "CSS selector for the repeating item, used to detect when clicking stopped adding content.",
      +  "type": "string"
      +}
    • addedInput schema / properties / load_more_scroll
      Added value: +{
      +  "description": "Scroll between load_more clicks. Defaults to true.",
      +  "type": "boolean"
      +}
    • addedInput schema / properties / max_chars
      Added value: +{
      +  "description": "Truncate the Markdown to this many characters. Prefer `query` + `chunking` when you want the relevant part of a long page rather than an arbitrary prefix.",
      +  "minimum": 1,
      +  "type": "integer"
      +}
    • addedInput schema / properties / max_chunks
      Added value: +{
      +  "description": "Maximum chunks to return (1–2000). Defaults to 500.",
      +  "maximum": 2000,
      +  "minimum": 1,
      +  "type": "integer"
      +}
    • addedInput schema / properties / query
      Added value: +{
      +  "description": "Natural-language question. Returns only the most relevant chunks, ranked (BM25), instead of the whole page. Requires chunking (enabled automatically when set).",
      +  "maxLength": 512,
      +  "type": "string"
      +}
    • addedInput schema / properties / query_top_k
      Added value: +{
      +  "description": "How many ranked chunks to return when `query` is set. Defaults to 5.",
      +  "maximum": 25,
      +  "minimum": 1,
      +  "type": "integer"
      +}
    • addedOutput schema / properties / ai_safety
      Added value: +{
      +  "description": "Prompt-injection report; present when ai_sanitize is true"
      +}
    • addedOutput schema / properties / chunks
      Added value: +{
      +  "description": "Chunk objects, ranked by relevance when `query` is set"
      +}
    • addedOutput schema / properties / debug
      Added value: +{
      +  "description": "Whether rendering, a proxy, or retries were used"
      +}
    • addedOutput schema / properties / headings
      Added value: +{}
    • addedOutput schema / properties / images
      Added value: +{}
    • changedOutput schema / properties / length / description
      Previous value: -"Character count of the Markdown content"New value: +"Character count of the returned Markdown"
    • addedOutput schema / properties / links
      Added value: +{}
    • addedOutput schema / properties / metadata
      Added value: +{
      +  "description": "Title, description, language, canonical and final URL"
      +}
    • removedOutput schema / properties / requestInfo
      Removed value: -{}
    • addedOutput schema / properties / request_id
      Added value: +{
      +  "type": "string"
      +}
    • addedOutput schema / properties / usage
      Added value: +{
      +  "description": "Character/token counts and truncation status"
      +}
  2. Addedv1.3.6

TDQS

A5/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already mark this as read-only, idempotent, open-world, and non-destructive, and the description adds valuable behavior beyond that: it strips navigation/ads/boilerplate by default, caps the Markdown text block at 6,000 characters while keeping full content in the structured field, and discloses prompt-injection scanning and HTTP 422 blocking behavior. No contradiction exists between the description and annotations.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness5/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The description is long, but the tool has 40 parameters and multiple behavioral nuances, so the length is justified. It is well structured with clear section headers (LONG PAGES, EXTRAS, UNTRUSTED CONTENT, Pick the right tool), front-loads the core purpose, and uses bolded parameter names for scannability without wasted sentences.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a complex 40-parameter, 1-required-parameter tool with an output schema and rich annotations, the description covers everything an agent needs to invoke it correctly: purpose, sibling routing, chunking strategy, security handling, output cap behavior, and tuning knobs. The existence of an output schema means return-value details do not need to be repeated in the description.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters5/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the baseline is 3, but this description adds substantial cross-parameter meaning: it explains why query+chunking is preferred over max_chars, that include_links/include_images/include_headings avoid a second scrape call, how ai_sanitize_mode coordinates with ai_sanitize, and how include_tags/exclude_tags control page sections. This goes well beyond the individual schema descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description names a specific verb and resource: converts any URL's HTML into clean Markdown via the OpenGraph.io v3 markdown endpoint. It distinguishes itself from siblings by describing what the output contains (main-content prose, headings, links, images) and by including an explicit 'Pick the right tool' list that contrasts getOgMarkdown with getOgData, getOgScrapeData, getOgExtract, getOgScreenshot, and getOgQuery.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage guidance is explicit and actionable: it tells agents to prefer query+chunking over max_chars for long pages, use include_tags/exclude_tags to target sections, and set ai_sanitize when feeding untrusted content to a model. The 'Pick the right tool' section directly lists when each sibling should be used instead, so an agent can route correctly without inferring.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.