Skip to main content
Glama

срезAI — Search API for AI agents

Обойти сайт / Crawl a site

crawl
Read-onlyIdempotent

Читает разделы сайта и читает найденные страницы: сначала карта адресов (как в map), затем текст страниц. Один вызов вместо «нашёл ссылку — прочитал — нашёл следующую»: так собирают всё содержимое раздела, а не отдельные страницы.

Когда: нужно содержимое нескольких страниц одного сайта — «собери все события с этого агрегатора», «прочитай раздел документации», «что в этом каталоге». Когда не: нужен только список адресов — map (дешевле, страницы не читаются); нужна одна страница — read_url. Фильтр search у crawl сравнивается с ТЕКСТОМ прочитанных страниц (не с адресом), а selectPaths/excludePaths отбирают САМИ адреса. Порядок такой: сначала map (увидеть разделы и адреса), затем crawl с selectPaths по нужному разделу и, если надо, с search по тексту. Возвращает: pages (адрес, заголовок, текст до 20 000 символов на страницу) и skipped — адреса, которые не прочитались, с причиной. Недоступные страницы не роняют чтение: остальные отдаются, а причина названа. Цена: карта (1 кредит за 10 возвращённых адресов, минимум 1) плюс 1 кредит за каждую прочитанную страницу. Не прочиталась — не считается.

Crawls a site and reads the pages it finds: first the URL map (as in map), then the page text. One call instead of "found a link — read it — found the next one": this is how you collect a whole section rather than single pages.

Use when: you need the content of several pages of one site — "collect every event from this aggregator", "read this docs section". Do not use when: you only need the list of URLs — map (cheaper, no page reads); you need one page — read_url. The search filter here matches the TEXT of the pages read (not the URL): search: "conference" keeps pages that mention it. So the order is: map with a URL filter first (to see the structure), then crawl with a text filter. Returns: pages (URL, title, up to 20 000 chars of text each) and skipped — URLs that could not be read, with the reason. Unreachable pages do not break the crawl: the rest are returned and the reason is stated. Cost: the map (1 credit per 10 URLs returned, minimum 1) plus 1 credit per page read. A page that failed to read is not counted.

Input Schema

TableJSON Schema
NameRequiredDescriptionDefault
urlYesСайт или страница (http/https) / Site or page URL
limitNoСколько страниц прочитать (1–20, по умолчанию 5). / How many pages to read (1–20, default 5).
searchNoОставить страницы, в тексте или заголовке которых есть эти слова (фильтр по СОДЕРЖИМОМУ прочитанного). / Keep pages whose text or title contains these words (a filter over the CONTENT read).
selectPathsNoЧитать только страницы, адрес которых содержит любую из этих подстрок — это и есть прицельное чтение раздела. Без него crawl читает первые адреса карты, а у агрегаторов это служебные страницы, а не события. / Read only pages whose URL contains any of these substrings — this is how you target a section. Without it crawl reads the first URLs of the map, which on aggregators are service pages, not events.
excludePathsNoНе читать страницы, чей адрес содержит любую из этих подстрок. / Do not read pages whose URL contains any of these substrings.
includeSubdomainsNoСчитать своими и поддомены. / Treat subdomains as own.

Schema Changelog

Changes observed during successful MCP inspections.

  1. Changed1 schema field changed
    • changedInput schema / properties / selectPaths / description
      Previous value: -"Читать только страницы, адрес которых содержит любую из этих подстрок — это и есть прицельный обход раздела. Без него crawl читает первые адреса карты, а у агрегаторов это служебные страницы, а не события. / Read only pages whose URL contains any of these substrings — this is how you target a section. Without it crawl reads the first URLs of the map, which on aggregators are service pages, not events."New value: +"Читать только страницы, адрес которых содержит любую из этих подстрок — это и есть прицельное чтение раздела. Без него crawl читает первые адреса карты, а у агрегаторов это служебные страницы, а не события. / Read only pages whose URL contains any of these substrings — this is how you target a section. Without it crawl reads the first URLs of the map, which on aggregators are service pages, not events."
  2. Added

TDQS

A4.7/5.0
Behavior5/5

Does the description disclose side effects, auth requirements, rate limits, or destructive behavior?

Annotations already declare readOnly/idempotent/non-destructive/open-world, but the description adds substantial context they cannot: a precise cost model (1 credit per 10 returned URLs, 1 credit per page read, failed reads not charged), graceful degradation ('unreachable pages do not break the crawl'), a 20 000-char per-page cap, and a skipped-with-reason list. This is exactly the beyond-annotation disclosure expected.

Agents need to know what a tool does to the world before calling it. Descriptions should go beyond structured annotations to explain consequences.

Conciseness3/5

Is the description appropriately sized, front-loaded, and free of redundancy?

The content is well front-loaded (purpose, when/when-not, filters, returns, cost), but every sentence is duplicated in Russian and English, roughly doubling the length for no informational gain. The Russian block also adds a subtle semantic nuance about search vs selectPaths/URL filtering not fully mirrored in English, so the redundancy is not purely wasteful but still heavy.

Shorter descriptions cost fewer tokens and are easier for agents to parse. Every sentence should earn its place.

Completeness5/5

Given the tool's complexity, does the description cover enough for an agent to succeed on first attempt?

For a 6-parameter open-world reader with no output schema, the description supplies everything an agent needs: return shape (pages URL/title/text plus skipped with reasons), failure behavior, cost, and the interaction between the text filter and URL filters. Nothing material is left unstated.

Complex tools with many parameters or behaviors need more documentation. Simple tools need less. This dimension scales expectations accordingly.

Parameters4/5

Does the description clarify parameter syntax, constraints, interactions, or defaults beyond what the schema provides?

Schema coverage is 100%, so the baseline is 3, but the description earns more by clarifying semantics the schema only partially conveys: search matches page TEXT/title rather than URL, selectPaths/excludePaths match URL substrings, and the warning that omitting selectPaths makes crawl read the map's first (often service) URLs. That is meaningful disambiguation beyond the field descriptions.

Input schemas describe structure but not intent. Descriptions should explain non-obvious parameter relationships and valid value ranges.

Purpose5/5

Does the description clearly state what the tool does and how it differs from similar tools?

The opening states a specific verb and resource — crawl reads a site's URL map then the page text — and frames it as one call replacing the manual find-link/read loop. It explicitly distinguishes itself from siblings map and read_url, so an agent can select it without opening any schema.

Agents choose between tools based on descriptions. A clear purpose with a specific verb and resource helps agents select the right tool.

Usage Guidelines5/5

Does the description explain when to use this tool, when not to, or what alternatives exist?

There are dedicated 'Use when' and 'Do not use when' sections naming concrete triggers ('collect every event from this aggregator', 'read this docs section') and the alternatives that win in the excluded cases (map for URL lists, read_url for one page). It even prescribes the recommended call order (map first, then crawl with selectPaths/search).

Agents often have multiple tools that could apply. Explicit usage guidance like "use X instead of Y when Z" prevents misuse.

Try in Browser

Glama MCP Gateway

Add one secure layer between your agents and this server.

Resources