Обойти сайт / Crawl a site
crawlЧитает разделы сайта и читает найденные страницы: сначала карта адресов (как в map), затем текст страниц. Один вызов вместо «нашёл ссылку — прочитал — нашёл следующую»: так собирают всё содержимое раздела, а не отдельные страницы.
Когда: нужно содержимое нескольких страниц одного сайта — «собери все события с этого агрегатора», «прочитай раздел документации», «что в этом каталоге». Когда не: нужен только список адресов — map (дешевле, страницы не читаются); нужна одна страница — read_url. Фильтр search у crawl сравнивается с ТЕКСТОМ прочитанных страниц (не с адресом), а selectPaths/excludePaths отбирают САМИ адреса. Порядок такой: сначала map (увидеть разделы и адреса), затем crawl с selectPaths по нужному разделу и, если надо, с search по тексту. Возвращает: pages (адрес, заголовок, текст до 20 000 символов на страницу) и skipped — адреса, которые не прочитались, с причиной. Недоступные страницы не роняют чтение: остальные отдаются, а причина названа. Цена: карта (1 кредит за 10 возвращённых адресов, минимум 1) плюс 1 кредит за каждую прочитанную страницу. Не прочиталась — не считается.
Crawls a site and reads the pages it finds: first the URL map (as in map), then the page text. One call instead of "found a link — read it — found the next one": this is how you collect a whole section rather than single pages.
Use when: you need the content of several pages of one site — "collect every event from this aggregator", "read this docs section". Do not use when: you only need the list of URLs — map (cheaper, no page reads); you need one page — read_url. The search filter here matches the TEXT of the pages read (not the URL): search: "conference" keeps pages that mention it. So the order is: map with a URL filter first (to see the structure), then crawl with a text filter. Returns: pages (URL, title, up to 20 000 chars of text each) and skipped — URLs that could not be read, with the reason. Unreachable pages do not break the crawl: the rest are returned and the reason is stated. Cost: the map (1 credit per 10 URLs returned, minimum 1) plus 1 credit per page read. A page that failed to read is not counted.
Input Schema
| Name | Required | Description | Default |
|---|---|---|---|
| url | Yes | Сайт или страница (http/https) / Site or page URL | |
| limit | No | Сколько страниц прочитать (1–20, по умолчанию 5). / How many pages to read (1–20, default 5). | |
| search | No | Оставить страницы, в тексте или заголовке которых есть эти слова (фильтр по СОДЕРЖИМОМУ прочитанного). / Keep pages whose text or title contains these words (a filter over the CONTENT read). | |
| selectPaths | No | Читать только страницы, адрес которых содержит любую из этих подстрок — это и есть прицельное чтение раздела. Без него crawl читает первые адреса карты, а у агрегаторов это служебные страницы, а не события. / Read only pages whose URL contains any of these substrings — this is how you target a section. Without it crawl reads the first URLs of the map, which on aggregators are service pages, not events. | |
| excludePaths | No | Не читать страницы, чей адрес содержит любую из этих подстрок. / Do not read pages whose URL contains any of these substrings. | |
| includeSubdomains | No | Считать своими и поддомены. / Treat subdomains as own. |