List sitemap URLs
gvt_list_sitemap_urlsDiscovers and ranks URLs from a domain's XML sitemaps to find the most significant pages (homepages, product pages, recently updated content) before analyzing them. Supports URL-pattern filtering and automatically flags robots.txt-blocked and already-tested pages. Legal/policy boilerplate (privacy, terms, cookies, disclaimers, refunds) is flagged legalPage:true with a -25 significance penalty; use the legalPages filter to drop it (exclude) or isolate it for a policy-coverage audit (only). Parameters group into three jobs: discovery (url), ranking and paging (limit, sort), and filtering (include and exclude URL patterns, excludeDisallowed, excludeTested, excludeScheduled, legalPages).
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | The target website URL or an explicit sitemap .xml URL. | |
| sort | No | How to sort the discovered URLs. 'significance' uses SEO heuristics. | significance |
| limit | No | Maximum number of URLs to return (default 10). | |
| exclude | No | Wildcard patterns to reject (e.g., '*archive*'). | |
| include | No | Wildcard patterns to require (e.g., '*blog*'). Only '*' wildcards are supported. | |
| legalPages | No | Tri-state filter for legal/policy boilerplate pages (privacy, terms, cookies, disclaimers, refund policies, etc). "include" keeps them (default), "exclude" drops them, and "only" returns just the legal/policy pages — useful for auditing a site's policy coverage. | include |
| excludeTested | No | If true, omits URLs the user has already tested. | |
| excludeScheduled | No | If true, omits URLs currently in the testing queue. | |
| excludeDisallowed | No | If true, silently drops URLs that are blocked by robots.txt. |
Output Schema
| Name | Required | Description | Default |
|---|---|---|---|
| sort | No | The sort actually applied to the results | |
| urls | No | Discovered URLs, ranked per the sort setting | |
| limit | No | Effective result cap | |
| errors | No | Per-sitemap fetch or parse failures; non-empty means discovery was partial | |
| hasMore | No | True when the filtered set was cut short by the limit | |
| matched | No | URLs remaining after include/exclude and legalPages filters | |
| returned | No | URLs actually returned after the limit | |
| truncated | No | True if the underlying sitemap parser hit its 5000-URL cap during fetching. | |
| legalPages | No | The legal-page filter actually applied | |
| bulkUrlLimit | No | How many of these URLs the caller's tier permits in one batch-analyze call. | |
| requestedUrl | No | The URL or sitemap URL that was resolved | |
| sitemapCount | No | Number of sitemaps parsed | |
| robotsTxtFound | No | Whether a robots.txt was found and consulted | |
| legalPagesFound | No | How many of the URLs discovered across all sitemaps were detected as legal/policy boilerplate, counted BEFORE any filtering or limiting is applied. | |
| totalDiscovered | No | Raw URL count discovered before filtering | |
| resolvedSitemaps | No | Sitemap URLs actually parsed |