Карта сайта / Map a site
mapСобирает адреса страниц сайта: robots.txt → Sitemap:, затем /sitemap.xml и /sitemap_index.xml (включая вложенные карты), а если карт нет — ссылки со стартовой страницы. Возвращает найденные адреса, их общее число и разбивку по разделам: по ней видно, где на сайте что лежит.
Когда: известен сайт, а не страница — «собрать все события с агрегатора», «какие разделы есть на сайте», «дай мне список страниц для чтения». Когда не: нужен текст страниц — crawl (он читает найденное); нужен ответ по теме без привязки к сайту — web_search. Фильтр search сравнивает подстроку с АДРЕСОМ (страницы мы не читаем), поэтому для русских сайтов указывайте часть латинского пути: search: "organizers" вместо «организаторы». Один вызов читает до 8 карт сайта: если сайт отдаёт больше, в ответе это сказано, а разделы показывают, куда идти дальше. Возвращает: pages (до limit адресов), total (сколько нашлось всего), sections (разделы с числом адресов), source (sitemap или links) и outcome: ok, blocked (сайт отвечает 403 — это не «страниц нет»), unreachable, empty. Цена: 1 кредит за каждые 10 возвращённых адресов, минимум 1 (200 адресов — 20 кредитов). Сайт, который нас не пустил, не тарифицируется: результата нет.
Collects a site's page URLs: robots.txt → Sitemap:, then /sitemap.xml and /sitemap_index.xml (including nested maps), and the start page's links when there are no maps. Returns the URLs, the total count, and a breakdown by section, which shows what lives where.
Use when: you know the site, not the page — "collect every event from this aggregator", "what sections does this site have", "give me the URLs to walk". Do not use when: you need the page text — crawl (it reads what map found); you need an answer about a topic — web_search. The search filter matches against the URL (map does not read pages), so for Russian sites pass part of the Latin path: search: "organizers". One call reads up to 8 sitemaps: if the site offers more, the answer says so, and sections show where to go next. Returns: pages (up to limit URLs), total (all found), sections (sections with URL counts), source (sitemap or links) and outcome: ok, blocked (the site answers 403 — that is not "no pages"), unreachable, empty. Cost: 1 credit per 10 URLs returned, minimum 1 (200 URLs — 20 credits). A site that blocks us is not charged: there is no result.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Сайт или страница (http/https) / Site or page URL | |
| limit | No | Сколько адресов вернуть (1–200, по умолчанию 50). total показывает, сколько нашлось всего. / How many URLs to return (1–200, default 50). total shows how many were found. | |
| search | No | Оставить адреса, содержащие эту подстроку (регистр и «ё» не важны; слова через пробел ищутся по отдельности). Фильтр по адресу, не по тексту страницы. / Keep URLs containing this substring (case- and ё-insensitive; space-separated words are matched individually). It filters URLs, not page text. | |
| selectPaths | No | Оставить только адреса, содержащие любую из этих подстрок (по одной на элемент). Так целятся в раздел: selectPaths: ["organizers"] на агрегаторе событий оставит 3565 адресов организаторов вместо служебных страниц. / Keep only URLs containing any of these substrings (one per item). This is how you aim at a section. | |
| excludePaths | No | Выбросить адреса, содержащие любую из этих подстрок: архивы, теги, служебные разделы. / Drop URLs containing any of these substrings: archives, tags, service sections. | |
| includeSubdomains | No | Считать своими и поддомены (по умолчанию только тот же хост). / Treat subdomains as own (by default only the same host). |