sitemap_extractor
Extract URLs from robots.txt, sitemap indexes, .xml.gz and text sitemaps, returning lastmod, changefreq, priority, plus new/removed diffs between runs.
Instructions
Sitemap URL extractor that reads robots.txt, sitemap indexes, .xml.gz and plain-text sitemaps and returns one row per URL with lastmod, changefreq, priority — plus a new/removed diff between runs. Billed to your own Apify account: ~$0.0002 per result (Apify free-plan price, lower on paid plans).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| mode | No | Mode — Choose list to dump every URL found right now, or diff to compare this run with the previous one and label each URL new, unchanged or removed. Diff mode needs at least two runs of the same source to be useful. Options: list = list — every URL in the sitemap; diff = diff — new / removed / unchanged since the last run. | list |
| sources | Yes | Sitemaps or domains — Enter the sitemaps to read, or just the domains, e.g. https://apify.com or https://www.allbirds.com/sitemap.xml. For a bare domain the Actor reads the Sitemap: lines of /robots.txt and falls back to /sitemap.xml, /sitemap_index.xml, /sitemap-index.xml, /sitemap.xml.gz and /sitemap.txt. Sitemap indexes are followed into their child sitemaps automatically. Example: ["https://apify.com"]. | |
| urlFilter | No | URL filter (regex) — Optional JavaScript regular expression; only URLs matching it are kept, e.g. /products/ for a Shopify catalogue or \.pdf$ for documents. Leave empty to keep every URL. The pattern is matched against the full URL and is case-sensitive, so write [Pp]roducts when you need both cases. | |
| maxUrlsPerSource | No | Max URLs per source — Enter how many URLs to keep per source, e.g. 5000. Counted across all child sitemaps of an index, so a 200,000-URL e-commerce site stops as soon as the cap is reached. Each URL is one billed dataset row. |