website_to_markdown
Turn any website into clean Markdown for LLMs, RAG pipelines, and vector DBs by crawling pages and converting content without a headless browser.
Instructions
Content Crawler turns any site into clean Markdown per page for LLMs, RAG pipelines and vector DBs — no headless browser, $1 per 1,000 pages. Billed to your own Apify account: ~$0.001 per result (Apify free-plan price, lower on paid plans).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| maxDepth | No | Max link depth — Enter how many links deep to follow from a start URL, e.g. 3. Set 0 to crawl only the start URL(s). | |
| maxPages | No | Max pages — Enter the maximum number of pages to crawl and convert, e.g. 50. This is the billed unit ($1 per 1,000 pages) — the crawl stops exactly at this count. | |
| startUrls | Yes | Start URLs — Enter the page(s) to start crawling from, e.g. https://docs.apify.com/platform. The crawler follows links from here in breadth-first order. Example: [{"url":"https://docs.apify.com/platform"}]. | |
| useSitemap | No | Seed from sitemap.xml — Turn this on to also read /sitemap.xml on each start URL's domain and add its URLs (filtered by the settings above) to the crawl queue, up to Max pages. | |
| outputFormat | No | Output format — Choose what content to put in each row: Markdown only, plain text only, or both. Markdown preserves headings, lists, code blocks and tables. Options: markdown = Markdown only; text = Plain text only; both = Markdown and plain text. | markdown |
| respectRobots | No | Respect robots.txt — Turn this on to skip URLs disallowed by the site's robots.txt file (recommended and on by default). | |
| sameDomainOnly | No | Same domain only — Turn this on to only follow links on the same domain as the start URL (www. is treated as the same domain), and off to also follow links to other domains. | |
| removeSelectors | No | Extra CSS selectors to remove — Optional. Extra CSS selectors to strip before extracting content, e.g. .cookie-banner or #newsletter-signup, on top of the built-in nav/header/footer/aside removal. | |
| excludePathPatterns | No | Exclude URL patterns — Optional. Regular expressions tested against the full URL; a match is skipped, e.g. \.(png|jpe?g)$ to skip images. Defaults cover binary files and login/signup pages. | |
| includePathPrefixes | No | Include path prefixes — Optional. Only crawl URLs whose path starts with one of these prefixes, e.g. /docs. Leave empty to crawl every path allowed by the other settings. |