Skip to main content
Glama

website_to_markdown

Turn any website into clean Markdown for LLMs, RAG pipelines, and vector DBs by crawling pages and converting content without a headless browser.

Instructions

Content Crawler turns any site into clean Markdown per page for LLMs, RAG pipelines and vector DBs — no headless browser, $1 per 1,000 pages. Billed to your own Apify account: ~$0.001 per result (Apify free-plan price, lower on paid plans).

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
maxDepthNoMax link depth — Enter how many links deep to follow from a start URL, e.g. 3. Set 0 to crawl only the start URL(s).
maxPagesNoMax pages — Enter the maximum number of pages to crawl and convert, e.g. 50. This is the billed unit ($1 per 1,000 pages) — the crawl stops exactly at this count.
startUrlsYesStart URLs — Enter the page(s) to start crawling from, e.g. https://docs.apify.com/platform. The crawler follows links from here in breadth-first order. Example: [{"url":"https://docs.apify.com/platform"}].
useSitemapNoSeed from sitemap.xml — Turn this on to also read /sitemap.xml on each start URL's domain and add its URLs (filtered by the settings above) to the crawl queue, up to Max pages.
outputFormatNoOutput format — Choose what content to put in each row: Markdown only, plain text only, or both. Markdown preserves headings, lists, code blocks and tables. Options: markdown = Markdown only; text = Plain text only; both = Markdown and plain text.markdown
respectRobotsNoRespect robots.txt — Turn this on to skip URLs disallowed by the site's robots.txt file (recommended and on by default).
sameDomainOnlyNoSame domain only — Turn this on to only follow links on the same domain as the start URL (www. is treated as the same domain), and off to also follow links to other domains.
removeSelectorsNoExtra CSS selectors to remove — Optional. Extra CSS selectors to strip before extracting content, e.g. .cookie-banner or #newsletter-signup, on top of the built-in nav/header/footer/aside removal.
excludePathPatternsNoExclude URL patterns — Optional. Regular expressions tested against the full URL; a match is skipped, e.g. \.(png|jpe?g)$ to skip images. Defaults cover binary files and login/signup pages.
includePathPrefixesNoInclude path prefixes — Optional. Only crawl URLs whose path starts with one of these prefixes, e.g. /docs. Leave empty to crawl every path allowed by the other settings.

Schema Changelog

Changes observed during successful MCP inspections.

  1. First observedv0.1.0

TDQS

B3.3/5.0
Behavior3/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

With no annotations, the description carries the full behavioral burden. It usefully discloses the cost model ('$1 per 1,000 pages', billed to the caller's own Apify account) and the no-headless-browser characteristic, but says nothing about credentials required, crawl duration, throttling, or what happens on failure.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

Only two sentences, but both are promotional and lead with the brand name rather than the operation. The detailed pricing in sentence two duplicates the maxPages schema description, so part of the text does not earn its place.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness3/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 10-parameter crawl tool with no annotations and no output schema, the description gives the gist of the result ('clean Markdown per page') and the cost, but omits auth/account requirements, per-result row shape, and run mechanics. It is adequate but leaves real gaps given the tool's complexity.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters3/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema description coverage is 100%, so the schema already documents all ten parameters thoroughly, including defaults, bounds and examples. The description adds no parameter-level meaning beyond the schema — its pricing line duplicates what maxPages already states.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose4/5

Does the description clearly state what the tool does and how it differs from similar tools?

The description states a concrete verb and resource ('turns any site into clean Markdown per page'), so the core operation is clear despite the marketing framing around 'Content Crawler'. It does not differentiate the tool from siblings like article_extractor or structured_data_extractor, which also extract content from pages.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines3/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

Usage is only implied via the audience callout ('for LLMs, RAG pipelines and vector DBs'), which tells the agent the broad context but not when to prefer this over article_extractor or structured_data_extractor. No explicit when-not conditions or alternative routing are provided.

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.